• Home
  • Search
  • A Simplified Framework for Using Multiple Imputation in Social Work Research
  • Cite Icon77
  • https://doi.org/10.1093/swr/32.3.171Copy DOI Icon

A Simplified Framework for Using Multiple Imputation in Social Work Research

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Missing data are nearly always a problem in research, and missing values represent a serious threat to the validity of inferences drawn from findings. Increasingly, social science researchers are turning to multiple imputation to handle missing data. Multiple imputation, in which missing values are replaced by values repeatedly drawn from conditional probability distributions, is an appropriate method for handling missing data when values are not missing completely at random. However, use of this method requires developing an imputation model from the observed data. This is typically a rigorous and time-consuming process. To encourage wider adoption of multiple imputation in social work research, a simple framework for designing imputation models is presented. The framework and its ability to generate unbiased estimates are demonstrated in a simulation study. KEY WORDS: missing data; multiple imputation; nonresponse ********** Missing data are ubiquitous in social research, and missingness or nonresponse can represent a threat to the validity of inferences because of undue effects on efficiency, power, and parameter bias (Shadish, Cook, & Campbell, 2002). Social work researchers are now addressing missing data in a more rigorous manner. Recently, Saunders et al. (2006) and Choi, Golder, Gillmore, and Morrison (2005) described important data imputation methods and dispelled misunderstandings regarding popular imputation methods, such as mean substitution. Recent advances in analytic methods, such as multiple imputation (MI), are taking hold in social work research. With MI, missing values are replaced with values repeatedly drawn from simulated conditional probability distributions (Schafer, 1997), thus creating multiple versions of the data set. Each version of the data set is analyzed according to the data analysis model, and the multiple results are combined into point estimates (Rubin, 1996). A critical task in MI is to devise an imputation model (Allison, 2002) or missing data model (Graham, Olchowski, & Gilreath, 2007), which involves specifying the measures that are putatively associated with the missing values. Although this process adds additional steps, the specification of an imputation model and the creation of multiple data sets can produce less-biased estimates in the presence of missing data across a wide variety of data analysis techniques (Schafer, 1997). Besides MI, there are many other methods for addressing missing data (Schafer, 1999; Schafer & Graham, 2002). An equally rigorous method known as direct or full information maximum likelihood (FIML) estimation can produce unbiased estimates and correct standard errors in the presence of missing data. When the number of imputations is sufficiently large, identical missing data models will produce the same estimates under MI and FIML (Graham et al., 2007). Unlike MI, FIML is limited to maximum likelihood analytic techniques and the missing data model must be included in the analysis model. Although we focus on MI, the steps we describe for developing an imputation model are equally appropriate for use in the missing data model for FIML. On the basis of the MI literature, this article describes a framework for developing an imputation model for use with any free or commercial software package that performs MI. BRIEF REVIEW OF MISSING DATA CONCEPTS Generally, both MI and a broad range of missing data issues have received ample attention in the applied literature (Graham et al., 2007). We briefly discuss the three types of distributions that describe the randomness of nonresponse given that this property has consequences for the development of an imputation model. For a discussion of general missing data concepts that are not critical to understanding our discussion of imputation model development, we refer readers to Schafer and Graham (2002). Distribution of Nonresponse The probability distribution of nonresponse--more frequently referred to as the missing data or nonresponse mechanism (Rubin, 1976)--is both an important factor in the decision to impute with MI and a context for the development of an imputation model. …

Similar Papers
  • Research Article
  • Citations56

Selecting the model for multiple imputation of missing data: Just use an IC!

  • Feb 24, 2021
  • Statistics in Medicine
  • Firouzeh Noghrehchi +3
  • Research Article
  • Citations20

New computations for RMSEA and CFI following FIML and TS estimation with missing data.

  • Apr 01, 2023
  • Psychological Methods
  • Xijuan Zhang +1
  • Research Article
  • Citations1

A Comparison of Multiple Imputation and Optimal Estimation for Missing and Uncertain Urban Air Toxics Data

  • Nov 01, 2006
  • Epidemiology
  • H Le +6
  • PDF
  • Research Article
  • Citations171

Outcome-sensitive multiple imputation: a simulation study

  • Jan 09, 2017
  • BMC Medical Research Methodology
  • Evangelos Kontopantelis +3
  • Research Article
  • Citations96

Multiple imputation in the presence of high-dimensional data

  • Jul 11, 2016
  • Statistical Methods in Medical Research
  • Yize Zhao +1
  • Research Article
  • Citations75

On Obtaining Estimates of the Fraction of Missing Information From Full Information Maximum Likelihood

  • Jul 20, 2012
  • Structural Equation Modeling: A Multidisciplinary Journal
  • Victoria Savalei +1
  • Research Article
  • Citations127

9. Multiple Imputation of Incomplete Categorical Data Using Latent Class Analysis

  • Jul 01, 2008
  • Sociological Methodology
  • Jeroen K Vermunt +3
  • Research Article
  • Citations1

The Life and Career of Matthew O. Howard

  • Mar 01, 2019
  • Journal of the Society for Social Work and Research
  • Jeffrey M Jenson
  • PDF
  • Research Article
  • Citations151

A comparison of different methods to handle missing data in the context of propensity score analysis

  • Oct 19, 2018
  • European Journal of Epidemiology
  • Jungyeon Choi +2
  • PDF
  • Research Article
  • Citations15

How to deal with missing longitudinal data in cost of illness analysis in Alzheimer's disease-suggestions from the GERAS observational study.

  • Jul 18, 2016
  • BMC Medical Research Methodology
  • Mark Belger +10
  • Supplementary Content

Pattern-mixture sensitivity analysis in longitudinaltrials with drop-out

  • Feb 26, 2016
  • LSHTM Research Online (London School of Hygiene and Tropical Medicine)
  • George Vamvakas
  • PDF
  • Research Article
  • Citations36

Multiple imputation with missing indicators as proxies for unmeasured variables: simulation study

  • Jul 08, 2020
  • BMC Medical Research Methodology
  • Matthew Sperrin +1
  • Research Article
  • Citations46

Multiple imputation analysis of case–cohort studies

  • Feb 24, 2011
  • Statistics in Medicine
  • Helena Marti +1
  • Research Article
  • Citations1005

The proportion of missing data should not be used to guide decisions on multiple imputation

  • Mar 13, 2019
  • Journal of Clinical Epidemiology
  • Paul Madley-Dowd +3
  • PDF
  • Research Article
  • Citations21

Should multiple imputation be stratified by exposure group when estimating causal effects via outcome regression in observational studies?

  • Feb 16, 2023
  • BMC Medical Research Methodology
  • Jiaxin Zhang +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.