• Home
  • Search
  • Predicting Missing Values in Survey Data Using Prompt Engineering for Addressing Item Non-Response
  • Cite Icon2
  • https://doi.org/10.3390/fi16100351Copy DOI Icon

Predicting Missing Values in Survey Data Using Prompt Engineering for Addressing Item Non-Response

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Survey data play a crucial role in various research fields, including economics, education, and healthcare, by providing insights into human behavior and opinions. However, item non-response, where respondents fail to answer specific questions, presents a significant challenge by creating incomplete datasets that undermine data integrity and can hinder or even prevent accurate analysis. Traditional methods for addressing missing data, such as statistical imputation techniques and deep learning models, often fall short when dealing with the rich linguistic content of survey data. These approaches are also hampered by high time complexity for training and the need for extensive preprocessing or feature selection. In this paper, we introduce an approach that leverages Large Language Models (LLMs) through prompt engineering for predicting item non-responses in survey data. Our method combines the strengths of both traditional imputation techniques and deep learning methods with the advanced linguistic understanding of LLMs. By integrating respondent similarities, question relevance, and linguistic semantics, our approach enhances the accuracy and comprehensiveness of survey data analysis. The proposed method bypasses the need for complex preprocessing and additional training, making it adaptable, scalable, and capable of generating explainable predictions in natural language. We evaluated the effectiveness of our LLM-based approach through a series of experiments, demonstrating its competitive performance against established methods such as Multivariate Imputation by Chained Equations (MICE), MissForest, and deep learning models like TabTransformer. The results show that our approach not only matches but, in some cases, exceeds the performance of these methods while significantly reducing the time required for data processing.

Similar Papers
  • Research Article
  • Citations10

Impact of Data Pre-Processing Techniques on XGBoost Model Performance for Predicting All-Cause Readmission and Mortality Among Patients with Heart Failure

  • Nov 01, 2024
  • BioMedInformatics
  • Qisthi Alhazmi Hidayaturrohman +1
  • Research Article

Machine learning-based fatigue classification using heart rate variability and cortisol: A multimodal approach to wearable health monitoring

  • May 01, 2025
  • Digital Health
  • Joung Eun Kim +5
  • Research Article

A Context-Aware Progressive Approach to Imputing Multivariate Heterogeneous Data in Water Pipe Networks

  • Jul 01, 2026
  • Journal of Computing in Civil Engineering
  • Hojat Behrooz +1
  • Research Article

Imputing PIRLS’s home socioeconomic status from students’ and schools’ questionnaire data: a simulation with MICE

  • Nov 25, 2025
  • International Journal of Testing
  • João Marôco +1
  • Book Chapter
  • Citations2

Missing Data Imputation for Continuous Variables Based on Multivariate Adaptive Regression Splines

  • Jan 01, 2020
  • Fernando Sánchez Lasheras +12
  • Research Article

Reconstructing aquifer dynamics with machine learning: Linking irrigation expansion to groundwater decline in a data-scarce hyper-arid region

  • Dec 01, 2025
  • Agricultural Water Management
  • Samuel Chucuya +11
  • PDF
  • Research Article
  • Citations156

Missing value imputation in high-dimensional phenomic data: imputable or not, and how?

  • Nov 05, 2014
  • BMC Bioinformatics
  • Serena G Liao +7
  • PDF
  • Research Article
  • Citations12

Comparison of Selected Multiple Imputation Methods for Continuous Variables – Preliminary Simulation Study Results

  • Feb 13, 2019
  • Acta Universitatis Lodziensis. Folia Oeconomica
  • Małgorzata Aleksandra Misztal
  • PDF
  • Research Article
  • Citations21

Study on the Missing Data Mechanisms and Imputation Methods

  • Jan 01, 2021
  • Open Journal of Statistics
  • Abdullah Z Alruhaymi +1
  • Book Chapter
  • Citations2

Random Forest Missing Data Imputation Methods: Implications for Predicting At-Risk Students

  • Aug 15, 2020
  • Bevan I Smith +2
  • PDF
  • Research Article
  • Citations17

High-Dimensional, Small-Sample Product Quality Prediction Method Based on MIC-Stacking Ensemble Learning

  • Dec 21, 2021
  • Applied Sciences
  • Jiahao Yu +2
  • Research Article
  • Citations6

LINE-1 methylation mediates the inverse association between body mass index and breast cancer risk: A pilot study in the Lebanese population

  • Apr 09, 2021
  • Environmental Research
  • Zainab Awada +10
  • Research Article
  • Citations5

Efficacy and Safety of Sovateltide in Patients with Acute Cerebral Ischaemic Stroke: A Randomised, Double-Blind, Placebo-Controlled, Multicentre, Phase III Clinical Trial

  • Jan 01, 2024
  • Drugs
  • Anil Gulati +18
  • Research Article
  • Citations2

Advancing Software Project Effort Estimation: Leveraging a NIVIM for Enhanced Preprocessing

  • Dec 18, 2024
  • Journal of Software: Evolution and Process
  • Syed Sarmad Ali +4
  • Research Article

Compqual-Tgnet: A Novel Hybrid Temporal-Graph Neural Architecture for Analyzing Competency and Quality Metrics in Oil and Gas Operations

  • Feb 07, 2025
  • International Journal For Multidisciplinary Research
  • Shashank Sawant -
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.