• Home
  • Search
  • Improving covariance-regularized discriminant analysis for EHR-based predictive analytics of diseases
  • Open Access IconOpen Access
  • Cite Icon6
  • https://doi.org/10.1007/s10489-020-01810-4Copy DOI Icon

Improving covariance-regularized discriminant analysis for EHR-based predictive analytics of diseases

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Linear Discriminant Analysis (LDA) is a well-known technique for feature extraction and dimension reduction. The performance of classical LDA however, significantly degrades on the High Dimension Low Sample Size (HDLSS) data for the ill-posed inverse problem. Existing approaches for HDLSS data classification typically assume the data in question are with Gaussian distribution and deal the HDLSS classification problem with regularization. However, these assumptions are too strict to hold in many emerging real-life applications, such as enabling personalized predictive analysis using Electronic Health Records (EHRs) data collected from an extremely limited number of patients who have been diagnosed with or without the target disease for prediction. In this paper, we revised the problem of predictive analysis of disease using personal EHR data and LDA classifier. To fill the gap, in this paper, we first studied an analytical model that understands the accuracy of LDA for classifying data with arbitrary distribution. The model gives a theoretical upper bound of LDA error rate that is controlled by two factors: (1) the statistical convergence rate of (inverse) covariance matrix estimators and (2) the divergence of the training/testing datasets to fitted distributions. To this end, we could lower the error rate by balancing the two factors for better classification performance. Hereby, we further proposed a novel LDA classifier De-Sparse that leverages De-sparsified Graphical Lasso to improve the estimation of LDA, which outperforms state-of-the-art LDA approaches developed for HDLSS data. Such advances and effectiveness are further demonstrated by both theoretical analysis and extensive experiments on EHR datasets https://www.overleaf.com/project/5d2728c718f6ff3b2bcf5991 .

Similar Papers
  • Research Article
  • Citations18

Daehr

  • Feb 08, 2017
  • ACM Transactions on Intelligent Systems and Technology
  • Haoyi Xiong +4
  • Research Article

AB1206 SELF-REPORT OF FRACTURE HISTORY COMPARED TO FRACTURE CODES FROM AN ELECTRONIC HEALTH RECORD DATASET

  • Jun 01, 2019
  • Annals of the Rheumatic Diseases
  • Maria Danila +5
  • Research Article

Modern approaches for assessing hospitalization risk and predictors based on health information system data. A systematic review

  • Feb 08, 2026
  • Cardiovascular Therapy and Prevention
  • R N Shepel +6
  • Research Article
  • Citations1

Performance of the FIND-FH machine learning algorithm for the identification of individuals with suspected familial hypercholesterolemia.

  • Aug 01, 2025
  • Journal of clinical lipidology
  • Spencer V Carter +11
  • PDF
  • Research Article
  • Citations3

Assessment of Projection Pursuit Index for Classifying High Dimension Low Sample Size Data in R

  • Jan 01, 2023
  • Journal of Data Science
  • Zhaoxing Wu +1
  • Research Article
  • Citations1

The Impact of Evolving Endometriosis Guidelines on Diagnosis and Observational Health Studies.

  • Dec 16, 2024
  • medRxiv : the preprint server for health sciences
  • Harry Reyes Nieva +13
  • PDF
  • Research Article
  • Citations9

Continuity and Completeness of Electronic Health Record Data for Patients Treated With Oral Hypoglycemic Agents: Findings From Healthcare Delivery Systems in Taiwan.

  • Apr 04, 2022
  • Frontiers in Pharmacology
  • Chien-Ning Hsu +7
  • Conference Article
  • Citations4

A novel method for handling Missing Not at Random Data in the electronic health records

  • Jun 01, 2022
  • Xinpeng Shen +4
  • Research Article
  • Citations11

Creation of a linked cohort of children and their parents in a large, national electronic health record dataset.

  • Aug 13, 2021
  • Medicine
  • Heather Angier +6
  • Conference Article
  • Citations13

Automated Prediction of Heart Disease Patients using Sparse Discriminant Analysis

  • Feb 01, 2019
  • K M Zubair Hasan +3
  • Conference Article
  • Citations1

Towards A Task Taxonomy of Visual Analysis of Electronic Health or Medical Record Data

  • Dec 01, 2018
  • Xiaohui Bian +5
  • PDF
  • Research Article
  • Citations18

A Computer Vision System for the Automatic Classification of Five Varieties of Tree Leaf Images

  • Jan 28, 2020
  • Computers
  • Sajad Sabzi +2
  • Research Article

The Lifecycle of Electronic Health Record Data in HIV-Related Big Data Studies: Qualitative Study of Bias Instances and Potential Opportunities for Minimization.

  • Aug 07, 2025
  • Journal of medical Internet research
  • Arielle N'Diaye +6
  • Abstract

OTHR-10. PILOT STUDY TO EVALUATE AUTOMATED LABORATORY DATA TRANSFER INTO MEDIDATA RAVE ACROSS THE NATIONAL CANCER INSTITUTE (NCI)-SUPPORTED PEDIATRIC BRAIN TUMOR CONSORTIUM (PBTC) AND PEDIATRIC EARLY PHASE CLINICAL TRIAL NETWORK (PEP-CTN) COOPERATIVE GROUPS

  • Jun 18, 2024
  • Neuro-Oncology
  • Tamara P Miler +13
  • Research Article
  • Citations19

Improved Sparse Multi-Class SVM and Its Application for Gene Selection in Cancer Classification

  • Jan 01, 2013
  • Cancer Informatics
  • Lingkang Huang +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.