• Home
  • Search
  • Data Extraction from Web Tables: The Devil is in the Details
  • Cite Icon23
  • https://doi.org/10.1109/icdar.2011.57Copy DOI Icon

Data Extraction from Web Tables: The Devil is in the Details

  • Sep 1, 2011
  • George Nagy +5 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We present a method based on header paths for efficient and complete extraction of labeled data from tables meant for humans. Although many table configurations yield to the proposed syntactic analysis, some require access to semantic knowledge. Clicking on one or two critical cells per table, through a simple interface, is sufficient to resolve most of these problem tables. Header paths, a purely syntactic representation of visual tables, can be transformed ("factored") into existing representations of structured data such as category trees, relational tables, and RDF triples. From a random sample of 200 web tables from ten large statistical web sites, we generated 376 relational tables and 34,110 subject-predicate-object RDF triples.

Similar Papers
  • Conference Article
  • Citations41

TCN: Table Convolutional Network for Web Table Interpretation

  • Apr 19, 2021
  • Daheng Wang +5
  • Book Chapter
  • Citations13

From Web Tables to Concepts: A Semantic Normalization Approach

  • Jan 01, 2015
  • Katrin Braunschweig +2
  • Research Article
  • Citations4

Data model for warehousing historical Web information

  • Mar 13, 2003
  • Information and Software Technology
  • Yinyan Cao +2
  • Conference Article
  • Citations1

IMAT: Intelligent Mobile Agent

  • Nov 01, 2018
  • Houssein Dhayne +2
  • Research Article
  • Citations3

Ekstraksi Data pada Tabel dari Halaman Web Menggunakan Pohon Document Object Model

  • Dec 27, 2016
  • Jurnal Nasional Teknik Elektro dan Teknologi Informasi (JNTETI)
  • Memen Akbar +2
  • Research Article
  • Citations20

Testing the Competition: Usability of Commercial Information Sites Compared with Academic Library Web Sites

  • Sep 01, 2002
  • College & Research Libraries
  • Tiffini Anne Travis +1
  • Conference Article
  • Citations6

MIDAS: Finding the Right Web Sources to Fill Knowledge Gaps

  • Apr 01, 2019
  • Xiaolan Wang +3
  • Research Article
  • Citations22

Deep web data extraction based on visual information processing

  • Oct 12, 2017
  • Journal of Ambient Intelligence and Humanized Computing
  • Jin Liu +4
  • Research Article
  • Citations4

Russian Web Tables: A Public Corpus of Web Tables for Russian Language Based on Wikipedia

  • Jan 01, 2023
  • Lobachevskii Journal of Mathematics
  • P E Fedorov +2
  • Conference Article
  • Citations3

The TextCap semantic interpreter

  • Jan 01, 2008
  • Charles B Callaway
  • Book Chapter
  • Citations1

Quantum Implementation of Database Operators and Queries

  • Jan 01, 2022
  • David Song +1
  • Research Article
  • Citations18

Artificial intelligence meets dairy cow research: Large language model's application in extracting daily time-activity budget data for a meta-analytical study.

  • Sep 01, 2025
  • Journal of dairy science
  • M Lamanna +6
  • Research Article
  • Citations12

Exploring the Feasibility of GPT-4 as a Data Extraction Tool for Renal Surgery Operative Notes.

  • May 31, 2024
  • Urology practice
  • Jessica Y Hsueh +6
  • Research Article

A pragmatic methodology to extract anesthetic and physiological data from the electronic health record.

  • Dec 06, 2023
  • Paediatric anaesthesia
  • Arshia Aalami Harandi +4
  • PDF
  • Research Article
  • Citations3

BioDB extractor: customized data extraction system for commonly used bioinformatics databases

  • Jun 01, 2015
  • BioData Mining
  • Rajiv Karbhal +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.