• Home
  • Search
  • FDTool: a Python application to mine for functional dependencies and candidate keys in tabular data.
  • Cite Icon4
  • https://doi.org/10.12688/f1000research.16483.1Copy DOI Icon

FDTool: a Python application to mine for functional dependencies and candidate keys in tabular data.

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Functional dependencies (FDs) and candidate keys are essential for table decomposition,database normalization, and data cleansing. In this paper, we present FDTool, a commandline Python application to discover minimal FDs in tabular datasets and infer equivalent attributesets and candidate keys from them. The runtime and memory costs associated withseven published FD discovery algorithms are given with an overview of their theoretical foundations.We conclude that FD_Mine is the most efficient FD discovery algorithm when appliedto datasets with many rows (> 100,000 rows) and few columns (< 14 columns). This putsit in a special position to rule mine clinical and demographic datasets, which often consistof long and narrow sets of participant records. The structure of FD Mine is described andsupplemented with a formal proof of the equivalence pruning method used. FDTool is are-implementation of FD Mine with additional features added to improve performance andautomate typical processes in database architecture. The experimental results of applyingFDTool to 12 datasets of different dimensions are summarized in terms of the number ofFDs checked, the number of FDs found, and the time it takes for the code to terminate. Wefind that the number of attributes in a dataset has a much greater effect on the runtime andmemory costs of FDTool than does row count. The last section explains in detail how theFDTool application can be accessed, executed, and further developed.

Loading PDF

Similar Papers
  • Research Article
  • Citations4

Preserving logical and functional dependencies in synthetic tabular data

  • Jul 01, 2025
  • Pattern Recognition
  • Chaithra Umesh +4
  • Conference Article
  • Citations7

Database normalization design pattern

  • Oct 01, 2017
  • Kunal Kumar +1
  • Research Article
  • Citations4

Improving Data Cleaning by Learning From Unstructured Textual Data

  • Jan 01, 2025
  • IEEE Access
  • Rihem Nasfi +2
  • Research Article
  • Citations95

Mining functional dependencies from data

  • Sep 15, 2007
  • Data Mining and Knowledge Discovery
  • Hong Yao +1
  • Conference Article
  • Citations23

Tabular Data Anomaly Patterns

  • Aug 01, 2017
  • Dina Sukhobok +2
  • Research Article
  • Citations7

Parallel Data Partitioning Algorithms for Optimization of Data-Parallel Applications on Modern Extreme-Scale Multicore Platforms for Performance and Energy

  • Jan 01, 2018
  • IEEE Access
  • Ravi Reddy Manumachu +1
  • Conference Article
  • Citations33

On the Discovery of Relaxed Functional Dependencies

  • Jan 01, 2016
  • Loredana Caruccio +2
  • Book Chapter

A Model for Analyzing and Visualizing Tabular Data

  • Jan 01, 2011
  • Ekaterina Simonenko +2
  • Book Chapter
  • Citations4

Mediated Data Integration Systems Using Functional Dependencies Embedded in Ontologies

  • Aug 20, 2011
  • Abdelghani Bakhtouchi +3
  • Book Chapter
  • Citations19

Revisiting Conditional Functional Dependency Discovery: Splitting the “C” from the “FD”

  • Jan 01, 2019
  • Joeri Rammelaere +1
  • PDF
  • Research Article
  • Citations4

A Self-Attention-Based Imputation Technique for Enhancing Tabular Data Quality

  • Jun 04, 2023
  • Data
  • Do-Hoon Lee +1
  • Research Article

Proposal of self and semi-supervised learning for imbalanced classification of coronary heart disease tabular data

  • Sep 09, 2024
  • Revista Tecnología en Marcha
  • Danny Xie-Li +1
  • Research Article
  • Citations33

An Efficient Algorithm to Compute the Candidate Keys of a Relational Database Schema

  • Feb 01, 1996
  • The Computer Journal
  • H Saiedian
  • PDF
  • Research Article
  • Citations4

SWITCHING-ALGEBRAIC ANALYSIS OF RELATIONAL DATABASES

  • Feb 01, 2014
  • Journal of Mathematics and Statistics
  • Rushdi
  • Research Article
  • Citations90

Efficient denial constraint discovery with hydra

  • Nov 01, 2017
  • Proceedings of the VLDB Endowment
  • Tobias Bleifuß +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.