• Home
  • Search
  • Towards Realistic Error Models for Tabular Data
  • Cite Icon3
  • https://doi.org/10.1145/3774914Copy DOI Icon

Towards Realistic Error Models for Tabular Data

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Errors in data are a key challenge in modern data management and processing systems. Monitoring and mitigating risks associated with errors in data transformations and downstream applications, such as Machine Learning (ML) model training, requires a profound understanding of error generation and impact of errors on data pipelines. Unfortunately, scientific progress in the field is facing two main challenges: For one, research on data errors often does not adhere to the FAIR (Findable, Accessible, Interoperable, and Reusable) principles, which impedes reproducibility and comparisons. Second, existing data error models are oversimplified and fail to capture the complex statistical dependencies underlying the types and distributions of errors observed in real-world data. Building on prior work in the database management systems and statistics literature, we extend the theory on missing values to encompass a broader range of errors in tables and provide an overview of relevant error types. Combining error sampling mechanisms often observed in real data with a comprehensive categorization of errors, we introduce a latent factor model for tabular data errors that is simple to implement and can effectively model realistic error dependencies. Error sampling is decoupled from error types, which allows for simple extensions with more error types or sampling mechanisms. Using established benchmarks, we evaluate our model in two application scenarios, data cleaning and tabular ML tasks. In a comprehensive suite of experiments we demonstrate the impact of realistic error models on data cleaning benchmarks. Our results also show that a simple generative error model captures a wide range of error mechanisms and offers a convenient formalization of data perturbations to improve the generalizability, robustness and reproducibility of data cleaning research.

Similar Papers
  • Research Article
  • Citations4

CALIFRAME: a proposed method of calibrating reporting guidelines with FAIR principles to foster reproducibility of AI research in medicine.

  • Oct 08, 2024
  • JAMIA open
  • Kirubel Biruk Shiferaw +4
  • Conference Article
  • Citations12

Intelli-Eye: An UAV Tracking System with Optimized Machine Learning Tasks Offloading

  • Apr 29, 2019
  • Bo Yang +6
  • Research Article

Accounting for correlated data errors in geomagnetic field modeling using Swarm magnetic observations

  • Nov 01, 2025
  • Physics of the Earth and Planetary Interiors
  • Clemens Kloss
  • Research Article
  • Citations206

A Sparsity-Driven Approach for Joint SAR Imaging and Phase Error Correction

  • Dec 09, 2011
  • IEEE Transactions on Image Processing
  • N Ö Onhon +1
  • Conference Article
  • Citations18

Design Space Exploration of Accelerators and End-to-End DNN Evaluation with TFLITE-SOC

  • Sep 01, 2020
  • Nicolas Bohm Agostini +6
  • Conference Article
  • Citations7

NetDiffusion: Network Data Augmentation Through Protocol-Constrained Traffic Generation

  • Jun 10, 2024
  • Xi Jiang +6
  • Research Article
  • Citations2

NetDiffusion: Network Data Augmentation Through Protocol-Constrained Traffic Generation

  • Jun 11, 2024
  • ACM SIGMETRICS Performance Evaluation Review
  • Xi Jiang +6
  • Conference Article
  • Citations6

On the Distribution of ML Workloads to the Network Edge and Beyond

  • May 10, 2021
  • Georgios Drainakis +4
  • Research Article
  • Citations13

HT2ML: An efficient hybrid framework for privacy-preserving Machine Learning using HE and TEE

  • Sep 27, 2023
  • Computers & Security
  • Qifan Wang +5
  • Preprint Article

Advances and prospects in hydrological (error) modelling

  • Mar 11, 2024
  • Carlo Albert
  • Research Article
  • Citations34

Limitations of machine learning for building energy prediction: ASHRAE Great Energy Predictor III Kaggle competition error analysis

  • Apr 21, 2022
  • Science and Technology for the Built Environment
  • Clayton Miller +3
  • Conference Article
  • Citations14

Towards Perspective-Based Specification of Machine Learning-Enabled Systems

  • Aug 01, 2022
  • Hugo Villamizar +2
  • PDF
  • Research Article
  • Citations77

Machine Learning in Dentistry: A Scoping Review.

  • Jan 25, 2023
  • Journal of clinical medicine
  • Lubaina T Arsiwala-Scheppach +4
  • PDF
  • Research Article

Delivering a machine learning course on HPC resources

  • Jan 01, 2020
  • EPJ Web of Conferences
  • Stefano Bagnasco +4
  • Research Article
  • Citations118

We need to talk about error: causes and types of error in veterinary practice

  • Oct 20, 2015
  • Veterinary Record
  • C Oxtoby +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.