- Book Chapter
2
- 10.1515/9783112208212-012
281Linguistic linked (open) data
- Dec 15, 2025
- Anas Fahad Khan
Publications from 2021 to 2026
Showing 10 of 68 papers
281Linguistic linked (open) data
Novel Benchmark for NER in the Wastewater and Stormwater Domain
Efficient wastewater and stormwater management is mandatory for sustainable cities. Extracting structured knowledge from reports and regulations is challenging due to domain-specific terminology and multilingual contexts. This work focuses on domain-specific Named Entity Recognition (NER) as a first step towards effective relation and information extraction to support decision making. A multilingual benchmark is crucial for evaluating these methods. This study develops a French-Italian domain-specific text corpus for wastewater management. It evaluates state-of-the-art NER methods, including LLM-based approaches, to provide a reliable baseline for future strategies and explores automated annotation projection in view of an extension of the corpus to new languages.
Read moreExploring the GDLI: a multidimensional approach to Historical Lexicography
This paper focuses on the challenges of creating a structured digital database from large datasets that were not originally digital and lack a rigorous organizational structure. A representative case is the computerization of Salvatore Battaglia's Grande Dizionario della Lingua Italiana (GDLI). This initiative part of a broader effort to recover, preserve, and enhance textual resources of historical, linguistic, and cultural significance. Due to its monumental nature, the original print format of the GDLI is highly complex. Although Optical Character Recognition (OCR) technology enables digitization, it is not immune to orthographic errors. The project, born as a collaboration between the Accademia della Crusca and the A. Zampolli Institute for Computational Linguistics (CNR-ILC), involves the creation of an online searchable database with advanced search capabilities. Currently digital dictionaries are widely available accessible on the Internet, attracting a diverse range of users. The GDLI, in particular, has been extensively consulted since becoming available online in its unstructured form. The project seeks to provide structured access tailored to the needs of linguists and language historians. The chosen work solutions and implementation strategies stem from a thorough analysis of the data and the potential uses of the dictionary’s information. The approach balances error management with the creation of a hybrid data representation model-one that is not a traditional database but rather a network of interconnected yet autonomous resources. The result is a multidimensional database capable of supporting diverse perspectives for analysis and consultation.
Read moreDiScEPT: a distributed environment for digital scholarly editions
The DiScEPT project aims to create an environment for the production and publication of digital editions. Its method-ology is grounded in multiple models that take into account document materiality, textual criticism, and, in particular, the analysis of parallel corpora. The project adopts a distributed environment that relies on a range of open-source software components, integrated through APIs. Each step of the digital edition process can thus be performed using highly specialized platforms that offer optimal solutions for specific tasks, minimizing the need for customization and ensuring interoperability. This paper presents the approach adopted for the implementation of the distributed environment, highlighting the role of CLARIN’s tools and services in supporting the development of the DiScEPT project and ensuring the long-term preservation of its data within a sustainable framework.
Read moreWhen Data Meets the Past: Data Collection, Sharing, and Reuse in Ancient World Studies
Abstract This article explores the challenges and opportunities of adopting data-driven approaches in Ancient World (AW) studies, focusing on the complexities of data collection, curation, and analysis in the field. We address issues such as defining data for AW studies, as well as data fragmentation, standardization, and interoperability. We propose solutions to enhance data accessibility, collaboration, and reuse, demonstrating that adopting standardized formats and adhering to FAIR principles can improve data sharing and enable large-scale, interdisciplinary research. Importantly, we highlight how qualitative and quantitative approaches can coexist, enriching the field. We also review different past and ongoing initiatives supporting data-driven methodologies in AW studies and advocate for their continued expansion. Lastly, we discuss the rise of data papers as a transformative tool for bridging traditional scholarship and digital methodologies, emphasizing the importance of data sets and their potential for reuse in advancing the field.
Read moreParlaMint II: advancing comparable parliamentary corpora across Europe
Abstract The paper presents the results of the ParlaMint II project, which comprise comparable corpora of parliamentary debates of 29 European countries and autonomous regions, covering at least the period from 2015 to 2022, and containing over 1 billion words. The corpora are uniformly encoded, contain rich metadata about their 24 thousand speakers, and are linguistically annotated up to the level of Universal Dependencies syntax and named entities. The paper focuses on the enhancement made since the ParlaMint I project and presents the compilation of the corpora, including the encoding infrastructure, use of GitHub, the production of individual corpora, the common pipeline for producing their distribution, and use of CLARIN services for dissemination. It then gives a quantitative overview of the produced corpora, followed by the qualitative additions made within the ParlaMint II project, namely metadata localisation, the addition of new metadata, such as the political orientation of political parties, the machine translation of the corpora to English and its tagging with semantic classes, and the production of pilot speech corpora. Finally, outreach activities and further work are discussed.
Read moreSub-word orthographic processing and semantic activation as revealed by ERPs
ABSTRACT The present study investigated ERP signatures of processing sub-word orthography, that is, shorter words embedded within longer words, in a non-priming task that pushes for semantics. Participants performed a semantic categorisation task on pseudosuffixed and nonsuffixed words, like corner and peace, that contained embedded words (corn, pea) either congruent or not with the probe category (e.g. FOOD vs. ANIMAL). While the task required semantic activation of the whole-word, sub-word orthography and semantics were activated. Results indicate stronger negativity from as early as ∼230 ms after word onset when the embedded word did not fit the category, but only for pseudosuffixed words. The observed neural dynamics point to rapid extraction of sub-word orthography and prompt activation of meaning thereupon. We discuss the results with respect to literature on morphological processing, context-dependency of the mechanisms, and interpretations of the N250 and N400 components.
Read moreExtracting Tuscan phonetic correspondences from dialect pronunciations automatically
Abstract We present a novel approach to identifying individual pairs of phonetic correspondences in a dataset of dialect pronunciations. This continues work identifying shibboleths (i.e., characteristic features of a given dialect), a category that has interested dialectology and that dialectometrical research has examined mostly in the form of categorical data or entire phonetic transcriptions. This article reaches into segmental sequences (phonetic transcriptions) to identify individual phonetic correspondences. We follow earlier work in examining how distinctive and how representative a given phonetic correspondence is for a selected group of varieties. We proceed from string alignments, and innovate in characterizing the important notions via information theory. Despite minor problems, the method improves on the generality of competing approaches and can be shown to be useful in detecting characteristic phonetic correspondences in Tuscan varieties. We argue that this facilitates deeper investigation into the relation between aggregating approaches to dialectology and approaches proceeding from features.
Read moreT-FREX: A Transformer-based Feature Extraction Method from Mobile App Reviews
Mobile app reviews are a large-scale data source for software-related knowledge generation activities, including software maintenance, evolution and feedback analysis. Effective extraction of features (i.e., functionalities or characteristics) from these reviews is key to support analysis on the acceptance of these features, identification of relevant new feature requests and prioritization of feature development, among others. Traditional methods focus on syntactic pattern-based approaches, typically context-agnostic, evaluated on a closed set of apps, difficult to replicate and limited to a reduced set and domain of apps. Meanwhile, the pervasiveness of Large Language Models (LLMs) based on the Transformer architecture in software engineering tasks lays the groundwork for empirical evaluation of the performance of these models to support feature extraction. In this study, we present T-FREX, a Transformer-based, fully automatic approach for mobile app review feature extraction. First, we collect a set of ground truth features from users in a real crowdsourced software recommendation platform and transfer them automatically into a dataset of app reviews. Then, we use this newly created dataset to fine-tune multiple LLMs on a named entity recognition task under different data configurations. We assess the performance of T-FREX with respect to this ground truth, and we complement our analysis by comparing T-FREX with a baseline method from the field. Finally, we assess the quality of new features predicted by T-FREX through an external human evaluation. Results show that T-FREX outperforms on average the traditional syntacticbased method, especially when discovering new features from a domain for which the model has been fine-tuned.
Read moreSub-word orthographic processing and semantic activation as revealed by ERPs
The present study investigated ERP signatures of processing sub-word orthography, that is, shorter words embedded within longer words, in a non-priming task that pushes for semantics. Participants performed a semantic categorization task on pseudosuffixed and nonsuffixed words, like corner and peace, that contained embedded words (corn, pea) either congruent or not with the probe category (e.g., FOOD vs. ANIMAL). While the task required semantic activation of the whole-word, sub-word orthography and semantics were activated. Results indicate stronger negativity from as early as ~230 ms after word onset when the embedded word did not fit the category, but only for pseudosuffixed words. The observed neural dynamics point to rapid extraction of sub-word orthography and prompt activation of meaning thereupon. We discuss the results with respect to literature on morphological processing, context-dependency of the mechanisms, and interpretations of the N250 and N400 components.
Read more