• Home
  • Search
  • Missing value imputation in high-dimensional phenomic data: imputable or not, and how?
  • Cite Icon156
  • https://doi.org/10.1186/s12859-014-0346-6Copy DOI Icon

Missing value imputation in high-dimensional phenomic data: imputable or not, and how?

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

BackgroundIn modern biomedical research of complex diseases, a large number of demographic and clinical variables, herein called phenomic data, are often collected and missing values (MVs) are inevitable in the data collection process. Since many downstream statistical and bioinformatics methods require complete data matrix, imputation is a common and practical solution. In high-throughput experiments such as microarray experiments, continuous intensities are measured and many mature missing value imputation methods have been developed and widely applied. Numerous methods for missing data imputation of microarray data have been developed. Large phenomic data, however, contain continuous, nominal, binary and ordinal data types, which void application of most methods. Though several methods have been developed in the past few years, not a single complete guideline is proposed with respect to phenomic missing data imputation.ResultsIn this paper, we investigated existing imputation methods for phenomic data, proposed a self-training selection (STS) scheme to select the best imputation method and provide a practical guideline for general applications. We introduced a novel concept of “imputability measure” (IM) to identify missing values that are fundamentally inadequate to impute. In addition, we also developed four variations of K-nearest-neighbor (KNN) methods and compared with two existing methods, multivariate imputation by chained equations (MICE) and missForest. The four variations are imputation by variables (KNN-V), by subjects (KNN-S), their weighted hybrid (KNN-H) and an adaptively weighted hybrid (KNN-A). We performed simulations and applied different imputation methods and the STS scheme to three lung disease phenomic datasets to evaluate the methods. An R package “phenomeImpute” is made publicly available.ConclusionsSimulations and applications to real datasets showed that MICE often did not perform well; KNN-A, KNN-H and random forest were among the top performers although no method universally performed the best. Imputation of missing values with low imputability measures increased imputation errors greatly and could potentially deteriorate downstream analyses. The STS scheme was accurate in selecting the optimal method by evaluating methods in a second layer of missingness simulation. All source files for the simulation and the real data analyses are available on the author’s publication website.Electronic supplementary materialThe online version of this article (doi:10.1186/s12859-014-0346-6) contains supplementary material, which is available to authorized users.

Loading PDF

Similar Papers
  • PDF
  • Research Article
  • Citations21

Study on the Missing Data Mechanisms and Imputation Methods

  • Jan 01, 2021
  • Open Journal of Statistics
  • Abdullah Z Alruhaymi +1
  • Research Article

A Context-Aware Progressive Approach to Imputing Multivariate Heterogeneous Data in Water Pipe Networks

  • Jul 01, 2026
  • Journal of Computing in Civil Engineering
  • Hojat Behrooz +1
  • PDF
  • Research Article
  • Citations12

Comparison of Selected Multiple Imputation Methods for Continuous Variables – Preliminary Simulation Study Results

  • Feb 13, 2019
  • Acta Universitatis Lodziensis. Folia Oeconomica
  • Małgorzata Aleksandra Misztal
  • Book Chapter
  • Citations2

Missing Data Imputation for Continuous Variables Based on Multivariate Adaptive Regression Splines

  • Jan 01, 2020
  • Fernando Sánchez Lasheras +12
  • Book Chapter
  • Citations2

Random Forest Missing Data Imputation Methods: Implications for Predicting At-Risk Students

  • Aug 15, 2020
  • Bevan I Smith +2
  • Research Article

Imputing PIRLS’s home socioeconomic status from students’ and schools’ questionnaire data: a simulation with MICE

  • Nov 25, 2025
  • International Journal of Testing
  • João Marôco +1
  • Research Article

Reconstructing aquifer dynamics with machine learning: Linking irrigation expansion to groundwater decline in a data-scarce hyper-arid region

  • Dec 01, 2025
  • Agricultural Water Management
  • Samuel Chucuya +11
  • Research Article
  • Citations180

Dealing with missing values in large-scale studies: microarray data imputation and beyond

  • Dec 04, 2009
  • Briefings in Bioinformatics
  • T Aittokallio
  • Research Article
  • Citations6

LINE-1 methylation mediates the inverse association between body mass index and breast cancer risk: A pilot study in the Lebanese population

  • Apr 09, 2021
  • Environmental Research
  • Zainab Awada +10
  • PDF
  • Research Article
  • Citations15

How to deal with missing longitudinal data in cost of illness analysis in Alzheimer's disease-suggestions from the GERAS observational study.

  • Jul 18, 2016
  • BMC Medical Research Methodology
  • Mark Belger +10
  • PDF
  • Research Article
  • Citations3

Classifying Incomplete Gene-Expression Data: Ensemble Learning with Non-Pre-Imputation Feature Filtering and Best-First Search Technique.

  • Oct 30, 2018
  • International Journal of Molecular Sciences
  • Yuanting Yan +5
  • PDF
  • Research Article
  • Citations4

A Self-Attention-Based Imputation Technique for Enhancing Tabular Data Quality

  • Jun 04, 2023
  • Data
  • Do-Hoon Lee +1
  • Research Article
  • Citations6

Improved clinical data imputation via classical and quantum determinantal point processes

  • May 09, 2024
  • eLife
  • Skander Kazdaghli +3
  • Research Article
  • Citations10

Impact of Data Pre-Processing Techniques on XGBoost Model Performance for Predicting All-Cause Readmission and Mortality Among Patients with Heart Failure

  • Nov 01, 2024
  • BioMedInformatics
  • Qisthi Alhazmi Hidayaturrohman +1
  • Research Article

Compqual-Tgnet: A Novel Hybrid Temporal-Graph Neural Architecture for Analyzing Competency and Quality Metrics in Oil and Gas Operations

  • Feb 07, 2025
  • International Journal For Multidisciplinary Research
  • Shashank Sawant -
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.