• Home
  • Search
  • How much missing data is too much? A single study exploration
  • Cite Icon1
  • https://doi.org/10.1200/jco.2006.24.18_suppl.6116Copy DOI Icon

How much missing data is too much? A single study exploration

Show More
  • Abstract
  • Literature Map
  • Citations
  • Similar Papers
Abstract

6116 Background: Analyses of patient-reported outcomes rely on the dependability of patients to complete and submit assessments in a timely manner – not all data is obtained. In recent work focusing on quality of life (QOL) data and imputation, it has been found that most methods do not alter study results. But how much data can be missing before study results are affected? Methods: Missing data was investigated using a 2-arm study (109 patients) who completed Linear Analogue Self Assessments at 4 intervals. Patients (11%) had missing data at the second interval. Existing data was analysed for differences in scores between arms, then cases were randomly deleted to create increasing percentages (12%-20%) of missing data. Ten simulations were conducted per percent. Imputation methods applied were carrying forward the last value (LVCF), average value (AVCF), and maximum value (MVCF). Student’s t-tests were performed between arms for each simulation. Results: Imputation did not alter results of our study data which was statistically significant (SS) between arms for overall QOL (p=0.036) and spiritual well-being (SWB) (p=0.006), and not statistically significant (NS) for mental well-being (MWB) (p=0.174). After data deletion and t-test calculations, AVCF did not impact results. For overall QOL, data deletion changed the p-value to NS in 1 of 10 simulations starting at 12% missing data and 5 of 10 simulations starting at 16% missing data. No matter what percentage of missing data, imputation produced a SS p-value over 80% of the time. Data deletion and subsequent imputation did not affect the study decision for SWB. For MWB, all differences between arms were NS prior to imputation. After imputation, there was at most a 7% disagreement in conclusions. LVCF and MVCF performed equally in all simulations. Conclusions: For this particular study, when p-values are close to the study-defined alpha, the increase in missing data can change the study results and imputation methods are more likely to determine SS differences. The further the p-values are from the study alpha, there is little effect from increasing missing data or applying imputation. These results are for one particular study and further research is needed. No significant financial relationships to disclose.

Similar Papers
  • Research Article
  • Citations10

Missing data techniques in classification for cardiovascular dysautonomias diagnosis.

  • Sep 24, 2020
  • Medical & Biological Engineering & Computing
  • Ali Idri +3
  • PDF
  • Research Article
  • Citations15

How to deal with missing longitudinal data in cost of illness analysis in Alzheimer's disease-suggestions from the GERAS observational study.

  • Jul 18, 2016
  • BMC Medical Research Methodology
  • Mark Belger +10
  • Research Article
  • Citations9

A comparison of missing data methods for hypothesis tests of the treatment effect in substance abuse clinical trials: a Monte-Carlo simulation study

  • Jun 03, 2008
  • Substance Abuse Treatment, Prevention, and Policy
  • Sarra L Hedden +2
  • PDF
  • Research Article
  • Citations178

Assessment of BSRN radiation records for the computation of monthly means

  • Feb 23, 2011
  • Atmospheric Measurement Techniques
  • A Roesch +5
  • Research Article
  • Citations82

Spiritual correlates of functional well-being in women with breast cancer.

  • Jun 01, 2002
  • Integrative Cancer Therapies
  • Ellen G Levine +1
  • Research Article
  • Citations284

Detecting changes in morphospace occupation patterns in the fossil record: characterization and analysis of measures of disparity

  • Jan 01, 2001
  • Paleobiology
  • Charles N Ciampaglio +2
  • Research Article

032 TRENDS IN HOSPITAL-LEVEL EFFECTS ATTRIBUTABLE TO MORTALITY AFTER ACUTE MYOCARDIAL INFARCTION: A STUDY OF 698 092 PATIENTS FROM THE MYOCARDIAL ISCHAEMIA NATIONAL AUDIT PROJECT (MINAP) 2004–2010

  • May 01, 2013
  • Heart
  • W R Long +1
  • Conference Article
  • Citations117

An evaluation of k-nearest neighbour imputation using likert data

  • Dec 23, 2004
  • P Jonsson +1
  • Front Matter
  • Citations8

Quality of Life and Depression in CKD: Improving Hope and Health

  • Aug 20, 2009
  • American Journal of Kidney Diseases
  • Suzanne Watnick
  • Research Article
  • Citations11

Road link traffic speed pattern mining in probe vehicle data via soft computing techniques

  • Jun 01, 2013
  • Applied Soft Computing
  • Dewang Chen +2
  • Front Matter
  • Citations106

Trial Quality in Nephrology: How Are We Measuring Up?

  • Aug 19, 2011
  • American Journal of Kidney Diseases
  • Suetonia C Palmer +2
  • Research Article
  • Citations11

Simulation study on missing data imputation methods for longitudinal data in cohort studies

  • Oct 10, 2021
  • Zhonghua liu xing bing xue za zhi = Zhonghua liuxingbingxue zazhi
  • Y M Li +5
  • Research Article
  • Citations503

Missing value estimation for DNA microarray gene expression data: local least squares imputation.

  • Aug 27, 2004
  • Bioinformatics
  • Hyunsoo Kim +2
  • Research Article

Estimating longitudinal change in latent variable means: a comparison of non-negative matrix factorization and other item non-response methods

  • Jul 14, 2022
  • Journal of Statistical Computation and Simulation
  • Olawale F Ayilara +4
  • Research Article

Abstract P5-04-11: Quality of life and psychosocial disparities between ethnically and religiously diverse population with advanced breast cancer

  • Jun 13, 2025
  • Clinical Cancer Research
  • Michal Braun +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.