• Home
  • Search
  • Diffusion Models for Tabular Data Imputation and Synthetic Data Generation
  • Cite Icon11
  • https://doi.org/10.1145/3742435Copy DOI Icon

Diffusion Models for Tabular Data Imputation and Synthetic Data Generation

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Data imputation and data generation have important applications across many domains where incomplete or missing data can hinder accurate analysis and decision-making. Diffusion models have emerged as powerful generative models capable of capturing complex data distributions across various data modalities such as image, audio, and time series. Recently, they have been also adapted to generate tabular data. In this article, we propose a diffusion model for tabular data that introduces three key enhancements: (1) a conditioning attention mechanism, (2) an encoder–decoder transformer as the denoising network, and (3) dynamic masking. The conditioning attention mechanism is designed to improve the model’s ability to capture the relationship between the condition and synthetic data. The transformer layers help model interactions within the condition (encoder) or synthetic data (decoder), while dynamic masking enables our model to efficiently handle both missing data imputation and synthetic data generation tasks within a unified framework. We conduct a comprehensive evaluation by comparing the performance of diffusion models with transformer conditioning against state-of-the-art techniques such as Variational Autoencoders, Generative Adversarial Networks, and Diffusion Models, on benchmark datasets. Our evaluation focuses on the assessment of the generated samples with respect to three important criteria, namely: (1) machine learning efficiency, (2) statistical similarity, and (3) privacy risk mitigation. For the task of data imputation, we consider the efficiency of the generated samples across different levels of missing features. The results demonstrate average superior machine learning efficiency and statistical accuracy compared to the baselines, while maintaining privacy risks at a comparable level, particularly showing increased performance in datasets with a large number of features. By conditioning the data generation on a desired target variable, the model can mitigate systemic biases, generate augmented datasets to address data imbalance issues, and improve data quality for subsequent analysis. This has significant implications for domains such as healthcare and finance, where accurate, unbiased, and privacy-preserving data are critical for informed decision-making and fair model outcomes.

Similar Papers
  • Research Article
  • Citations4

Preserving logical and functional dependencies in synthetic tabular data

  • Jul 01, 2025
  • Pattern Recognition
  • Chaithra Umesh +4
  • Book Chapter
  • Citations15

Generation of Synthetic Tabular Healthcare Data Using Generative Adversarial Networks

  • Jan 01, 2023
  • Alireza Hossein Zadeh Nik +3
  • Dissertation

Data-Centric AI: Tabular Data Synthesis with Deep Generative Models

  • Sep 13, 2024
  • Alex Xing Wang
  • Preprint Article

Novel Generative Adversarial Network Architectures for Generating image Data

  • Jun 19, 2024
  • Sanaz Mohammad Jafari
  • Conference Article
  • Citations10

GenEthos: A Synthetic Data Generation System With Bias Detection And Mitigation

  • Jun 23, 2022
  • Shubham Gujar +6
  • Research Article
  • Citations2

Innovative synthetic EHR data generation: diffusion models for enhanced privacy and clinical utility in multimorbidity clustering

  • Oct 06, 2025
  • Connection Science
  • Francis John Kita +2
  • Research Article

Optimizing clustering of electronic health tabular data: generative adversarial networks and Dirichlet process mixture models for advance healthcare analytics

  • May 26, 2025
  • IISE Transactions on Healthcare Systems Engineering
  • Francis John Kita +2
  • Research Article
  • Citations3

Collaborative Structure-Preserved Missing Data Imputation for Single-Cell RNA-Seq Clustering.

  • Sep 01, 2024
  • IEEE/ACM transactions on computational biology and bioinformatics
  • Hang Gao +4
  • Research Article
  • Citations9

TransFusion: Generating long, high fidelity time series using diffusion models with transformers

  • Jun 01, 2025
  • Machine Learning with Applications
  • Md Fahim Sikder +2
  • Conference Article
  • Citations8

Leveraging synthetic data for AI bias mitigation

  • Jun 14, 2023
  • Ajay M Patrikar +2
  • Research Article

Generative Adversarial Networks for Synthetic Data Generation in Diabetic Patient Research: Techniques, Applications, and Challenges.

  • Oct 03, 2025
  • Studies in health technology and informatics
  • Antonio García-Domínguez +6
  • Conference Article
  • Citations11

Categorical EHR Imputation with Generative Adversarial Nets

  • Jun 01, 2019
  • Yinchong Yang +3
  • Research Article
  • Citations8

Boosting EEG and ECG Classification with Synthetic Biophysical Data Generated via Generative Adversarial Networks

  • Nov 22, 2024
  • Applied Sciences
  • Archana Venugopal +1
  • Research Article
  • Citations10

Generation of synthetic data with low-dimensional features for condition monitoring utilizing Generative Adversarial Networks

  • Jan 01, 2022
  • Procedia Computer Science
  • Wagner Fabian +4
  • Conference Article

The Collapse of Data: Comparing the Resiliency of Different AI Models to Model Collapse Using AI-Generated Image Data

  • Dec 22, 2025
  • Kevin Audreylius +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.