- Research Article
- 10.1145/3800678
Automating the Extraction of Structured Data from Large Newspaper Corpora using Layout Analysis, OCR and Generative AI
- Mar 06, 2026
- Journal on Computing and Cultural Heritage
- Nikos Kontonasios + 2 more +2
Historical newspapers are invaluable resources that capture the cultural, social, economic, and political dimensions of their time, providing unique insights into past events and daily life. However, the potential of these archives remains underutilised due to the limitations of traditional research methods, which rely on laborious manual examination of individual pages or articles. Digital archives have begun to address this issue by enhancing accessibility, but the lack of effective tools for parsing and analysing the contents still restricts their utility for large-scale historical research. Optical Character Recognition (OCR) technology plays a foundational role in converting scanned images into searchable text. However, OCR alone cannot adequately handle the complexities of historical newspapers, which often feature degraded image quality, inconsistent typography, and intricate layouts. In parallel, the effective segmentation and ordering of newspaper contents is essential for generating well-structured and logically ordered data from specific newspaper sections. This paper focuses on the development of a pipeline that addresses these challenges, using the daily newspaper Le Sémaphore de Marseille as a case study (35,703 issues covering the period 1827 to 1944). The proposed system integrates layout analysis, OCR, and Generative AI-driven information extraction techniques to extract specific data elements from scanned images, namely ship arrivals data, and to store them in a machine-readable format for further analysis. This well-structured output enables historians to analyse historical trade patterns, economic trends, and regional interactions more efficiently and at a much larger scale. An evaluation of the key pipeline components demonstrates the effectiveness of the system, achieving an F1-score of 96% in paragraph segmentation and information extraction, while also highlighting the need for further improvement in the critical OCR component.
Read more