• Home
  • Search
  • Data-Centric AI: Tabular Data Synthesis with Deep Generative Models
  • https://doi.org/10.26686/wgtn.27014419Copy DOI Icon

Data-Centric AI: Tabular Data Synthesis with Deep Generative Models

  • Sep 13, 2024
  • Alex Xing Wang
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

<p><strong>Because of its wide range of applications, generative artificial intelligence (Generative AI) has received a lot of attention in academia and industry. Despite significant advances in computer vision and natural language processing, the use of Generative AI in tabular data is still underexplored. This gap is especially significant given the prevalence of tabular data as the primary data modality. This thesis seeks to close this gap by focusing on the efficient generation, evaluation, refinement, and application of tabular data, while addressing the challenges inherent in its heterogeneous nature.</strong></p><p>The primary goal of this thesis is to develop efficient algorithms for tabular data synthesis, advancing the field of Generative AI in the tabular domain. To achieve this goal, four specific research objectives have been outlined.</p><p>First, this thesis addresses gaps in evaluation metrics by unifying a framework for the comprehensive and consistent assessment of synthetic data. The proposed framework is designed to meet diverse business requirements across various downstream tasks by incorporating a wide range of advanced metrics. These metrics cover various types of evaluations, including univariate, bivariate, multivariate, cluster, and record-level evaluations. Additionally, standardized visualizations are provided to facilitate qualitative assessments. The results indicate that the proposed evaluation framework not only enables consistent comparisons, rankings, and the selection of different data synthesis approaches but also acts as a valuable tool to promptly assess the reliability of results obtained from synthetic data.</p><p>Second, this thesis presents innovative algorithms for the generation of tabular data, aiming to significantly enhance the quality of data synthesis. One of the contributions is the development of a reversible feature engineering pipeline designed to automatically represent tabular data in an efficient format while also ensuring that the transformed data can be easily converted back to its original format. Additionally, the thesis proposes novel deep learning-based tabular data generation models that are capable of learning the joint distribution of multivariate datasets without relying on predefined distribution assumptions. Experimental results highlight the effectiveness of the proposed algorithms in handling heterogeneous data, demonstrating superior performance in most scenarios within the same training duration when compared to similar alternatives.</p><p>Third, this thesis introduces innovative synthetic data prototype selection algorithms with the goal of refining the generated samples. This approach aims to leverage the advantages of Deep Generative Models, which, once trained, can generate unlimited and diverse synthetic data. By carefully identifying and selecting high-quality samples or removing unrealistic ones, the quality of the synthetic data can be enhanced from a post-processing perspective. Building on this hypothesis, we recognize the iterative nature inherent in the data synthesis procedure, where the processes of data generation, evaluation, and refinement should operate repeatedly in a cyclical flow. During the data generation phase, the evaluation framework assists in identifying potential issues and risks based on use cases, while the refinement (post-processing) step iteratively improves synthetic data in alignment with the evaluation outcomes.</p><p>Last, this thesis applies and validates the proposed models across various domains, including business, healthcare, and government, each requiring distinct downstream tasks. Specifically, we assess our algorithms from the standpoint of data balancing using churn data, evaluate our models focusing on data augmentation with health data, and test our algorithms from the perspective of data representation considering data privacy concerns within sensitive citizen data. These applications serve as practical illustrations, demonstrating the effectiveness and utility of the proposed Generative AI models in real-world scenarios.</p><p>In summary, this thesis makes a substantial contribution to the advancement of Generative AI within the domain of tabular data. It not only presents innovative algorithms and evaluation methods but also introduces practical frameworks applicable to real-world scenarios. The findings demonstrate the considerable capacity of Generative AI to fundamentally transform tabular data in various fields, ultimately leading to improved data availability, quality, and quantity.</p>

Similar Papers
  • Research Article
  • Citations4

Preserving logical and functional dependencies in synthetic tabular data

  • Jul 01, 2025
  • Pattern Recognition
  • Chaithra Umesh +4
  • Book Chapter
  • Citations15

Generation of Synthetic Tabular Healthcare Data Using Generative Adversarial Networks

  • Jan 01, 2023
  • Alireza Hossein Zadeh Nik +3
  • Research Article

Probabilistic Versus Deep Generative Models: A Fairness Centred Evaluation of Synthetic Healthcare Tabular Data

  • Feb 26, 2026
  • International Journal of Computational Intelligence Systems
  • Dima Alattal +4
  • PDF
  • Research Article
  • Citations289

Deep Generative Models in Engineering Design: A Review

  • Mar 18, 2022
  • Journal of Mechanical Design
  • Lyle Regenwetter +2
  • Research Article

Generative Artificial Intelligence in Urban Design: A Review of Recent Applications

  • Dec 01, 2025
  • Landscape Architecture
  • Qiyuan Hong +2
  • Conference Article

On Causal and Anticausal LLM-based Data Synthesis

  • Feb 16, 2026
  • Bohan Jiang +5
  • PDF
  • Research Article
  • Citations15

Generative AI for cyber threat intelligence: applications, challenges, and analysis of real-world case studies

  • Aug 20, 2025
  • Artificial Intelligence Review
  • Prasasthy Balasubramanian +6
  • Research Article
  • Citations1

A Framework for Generating Realistic Synthetic Tabular Data in a Randomized Controlled Trial Setting.

  • Aug 01, 2025
  • Statistics in medicine
  • Niki Z Petrakos +2
  • Research Article

Optimizing generative AI models for edge deployment: Techniques and best practices

  • Apr 30, 2025
  • World Journal of Advanced Research and Reviews
  • Sai Kalyan Reddy Pentaparthi
  • Supplementary Content

Cyberdefense Powered by Generative AI: A Comprehensive State-of-the-Art Review

  • Oct 24, 2025
  • Ahmed A Alsamman +1
  • Research Article
  • Citations1

Generative AI in Healthcare: An Analytical Review of Models, Clinical Applications, and Decision-Support Implications

  • Dec 31, 2025
  • Journal of Future Artificial Intelligence and Technologies
  • Naglaa Fadul +3
  • Research Article

Balanced Marginal and Joint Distributional Learning for Tabular Data Synthesis via Mixture Cramer–Wold Distance

  • Mar 18, 2026
  • Applied Sciences
  • Seunghwan An +2
  • Research Article
  • Citations8

Qualitatively different teacher experiences of teaching with generative artificial intelligence

  • May 26, 2025
  • International Journal of Educational Technology in Higher Education
  • Robert Ellis +2
  • Research Article
  • Citations11

Diffusion Models for Tabular Data Imputation and Synthetic Data Generation

  • Jul 21, 2025
  • ACM Transactions on Knowledge Discovery from Data
  • Mario Villaizán-Vallelado +3
  • PDF
  • Research Article
  • Citations5

Integrating Multimodal Generative AI Technologies in Postgraduate Marketing Education

  • Nov 23, 2024
  • ASCILITE Publications
  • Terrence Chong
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.