- Research Article
39
- 10.1016/j.culher.2021.11.002
Linked open data in authoring virtual exhibitions
- Jan 01, 2022
- Journal of Cultural Heritage
- Daniele Monaco + 3 more +3
Linked open data in authoring virtual exhibitions
In cultural heritage, many projects execute Named Entity Linking (NEL) through global Linked Open Data (LOD) references in order to identify and disambiguate entities in their local datasets. It allows users to obtain extra information and contextualise the data with it. Thus, the aggregation and integration of heterogeneous LOD are expected. However, such development is still limited partly due to data quality issues. In addition, analysis on the LOD quality has not sufficiently been conducted for cultural heritage. Moreover, most research on data quality concentrates on ontology and corpus level observations. This paper examines the quality of the eleven major LOD sources used for NEL in cultural heritage with an emphasis on instance-level connectivity and graph traversals. Standardised linking properties are inspected for 100 instances/entities in order to create traversal route maps. Other properties are also assessed for quantity and quality. The outcomes suggest that the LOD is not fully interconnected and centrally condensed; the quantity and quality are unbalanced. Therefore, they cast doubt on the possibility of automatically identifying, accessing, and integrating known and unknown datasets. This implies the need for LOD improvement, as well as the NEL strategies to maximise the data integration.
Linked open data in authoring virtual exhibitions
Linked open data in authoring virtual exhibitions
C3EL: A Joint Model for Cross-Document Co-Reference Resolution and Entity Linking
Cross-document co-reference resolution (CCR) computes equivalence classes over textual mentions denoting the same entity in a document corpus.Named-entity linking (NEL) disambiguates mentions onto entities present in a knowledge base (KB) or maps them to null if not present in the KB.Traditionally, CCR and NEL have been addressed separately.However, such approaches miss out on the mutual synergies if CCR and NEL were performed jointly.This paper proposes C3EL, an unsupervised framework combining CCR and NEL for jointly tackling both problems.C3EL incorporates results from the CCR stage into NEL, and vice versa: additional global context obtained from CCR improves the feature space and performance of NEL, while NEL in turn provides distant KB features for already disambiguated mentions to improve CCR.The CCR and NEL steps are interleaved in an iterative algorithm that focuses on the highest-confidence still unresolved mentions in each iteration.Experimental results on two different corpora, news-centric and web-centric, demonstrate significant gains over state-of-the-art baselines for both CCR and NEL.
Read moreOn-The-Fly Academic Linked Data Integration
The web of Linked Open Data (LOD) has a prominent and rapid evolution recently. Over the last few years, LOD had developed to involve a wide range of various domains. Due to these facts, and the great interconnections among linked open datasets, linked data integration task had gained a huge attention and became a focal point of research. LOD applications aim to incorporate data from different LOD sources. Unfortunately, these sources of data are heterogeneous in schema and/or in vocabularies. Due to this heterogeneity, numerous challenges are emerging that have to be overcome. In this paper, a LOD integration framework is proposed, which aims to tackle these challenges. It works on integrating academic LOD datasets that reside in different LOD repositories with intrinsic schema and vocabularies heterogeneity. An automatic mapping technique in the integration processes is proposed in this paper. Consequently, an obvious decrease in execution time for the entire integration process, as well as, a great progress in the integrated data quality assessment metrics has been achieved.
Read moreMetaBelgica Project
Trustworthy metadata about entities related to cultural heritage is important to correctly identify bibliographic meta-data, attribute works, identify works of public domain (based on contributors’ date of death), or to support Named Entity Linking (NEL) for digitised documents. However, in Belgium such data is currently dispersed between numerous institutions, modelled with different ontologies, represented in different formats and curated in different languages. GLAM institutions would profit from data about Belgian entities to correctly annotate their collections.Furthermore, this would facilitate research in Belgium and abroad. The current situation does not only complicate data exchange in a national and international setting, but also leads to duplicate data curation efforts and data quality issues, negatively impacting the users’ experience. This paper introduces the MetaBelgica project, coordinated by KBR - The Royal Library of Belgium with the Royal Museums of Fine Arts, the Royal institute for Cultural Heritage, and the Royal Museum of Art and History, to improve the status quo. The aforementioned Federal Scientific Institutes (FSIs), belonging to the Belgian Science Policy Office (BELSPO), joined forces to develop a shared Linked Data platform for managing shared entities. This platform, based on Wikibase (the technology behind Wikidata), aims to ensure FAIR data about persons, organisations, time/events and locations related to Belgian cultural heritage. We will integrate data about millions of Belgian entities from the four participating FSIs into a Wikibase instance and make it accessible by using Persistent Identifiers. This will not only professionalize our own data management and improve data quality, but according to received project support letters, also impact other regional, national and European institutions.
Read moreLOD for Library Science: Benefits of Applying Linked Open Data in the Digital Library Setting
Linked Open Data (LOD) has gained widespread adoption by large industries as well as non-profit organizations and governmental organizations. One of the early adopters of LOD technologies are libraries. Since the “early years”, libraries have been key use case and innovation driver for LOD and significantly contributed to the adoption of semantic technologies. The first part of this paper presents selected success stories of current activities in the Linked Data Library community. In a nutshell, these studies include (1) a conceptualization of the Linked Data Value chain, (2) a case study for consumption of Linked Data in a digital journal environment, and (3) an approach to publish metadata on the Semantic Web from an Open Access repository. These stories reveal a strong relationship between LOD in libraries and research topics addressed in traditional fields of computer science such as artificial intelligence, databases, and knowledge discovery. Thus, in the second part of this paper we systematically review the relation of LOD in digital libraries from a computer science perspective. We discuss current LOD research topics such as data integration and schema integration, distributed data management, and others. These challenges have been discussed with computer scientists at a German national database meetup as well as with librarians from ZBW—Leibniz Information Center for Economics and at international librarians meetup.
Read moreAutomatic Named Entity Linking for Ancient Greek with a Domain-Specific Knowledge Base
Named Entity Linking, or disambiguating named entities by linking them to a knowledge base, is an important Natural Language Processing task, especially in the humanities. In this paper, we examine the performance of the state-of-the-art entity linking model BLINK in connecting Ancient Greek person mentions to a domain-specific German-language knowledge base. To train the model, we create both gold-standard data through manual annotation and noisier silver data through automatic extraction. We then evaluate whether incorporating the latter improves performance. Our findings suggest that, overall, the results remain suboptimal for Ancient Greek. Increasing training data, even through automatic methods, shows promise. However, as it stands, using BLINK directly would be ill-suited for Named Entity Linking in the target setting. We discuss possible causes and suggest areas for improvement.
Read moreSemantic-Based Data Access to Digital Cultural Heritage: A Qualitative Analysis of Linked Open Data Implementation
Semantic-Based Data Access to Digital Cultural Heritage: A Qualitative Analysis of Linked Open Data Implementation
Visualizing the Information of a Linked Open Data Enabled Research Information System
The Open Access movement and the research management can take a new turn if the research information is published as Linked Open Data. With Linked Open Data, the management of the research information within institutions and across institutions can be facilitated, the quality of the available data can be improved and their availability to the public is assured. However, it can be difficult for non-expert users to take advantage of the interlinked information offered by Linked Open Data as they lack of in- depth knowledge. In this paper, we present a use case of publishing research metadata as Linked Open Data and creating interactive visualizations to support users in analyzing the Flemish research landscape.
Read moreOpen Data Integration for Lanna Cultural Heritage e-Museums
Cultural heritage is the way of life. Using digital technology to manage the information of cultural heritage has become an important issue on the perseverance of cultural heritage. Several cultural heritage institutions have their own large databases with different metadata schemas. Linked Open Data is a mechanism of publishing structured data that allows metadata to be connected. This paper presents a methodology for integrating, converting Lanna cultural heritage metadata and linking data to external data sources. Our framework uses OAI-PMH (Open Archives Initiative—Protocol for Metadata Harvesting) by converting the various metadata formats to Dublin Core metadata standard. The Dublin Core/XML was used for the integrated metadata to create a central repository of data sources. The OAI harvester was developed to extract data from the various repositories and gather data into the repository. A framework for publishing the repository data as a Linked Data source in the RDF format using the OAM Framework is described. In this framework, ontology is used as a common schema for publishing the RDF data from the database sources by mapping database schema to ontology. Finally, we demonstrate Linked Data consumption by creating a cultural heritage portal application prototype from the published RDF data sources.
Read moreIs dc:subject enough? A landscape on iconography and iconology statements of knowledge graphs in the semantic web
PurposeIn the last few years, the size of Linked Open Data (LOD) describing artworks, in general or domain-specific Knowledge Graphs (KGs), is gradually increasing. This provides (art-)historians and Cultural Heritage professionals with a wealth of information to explore. Specifically, structured data about iconographical and iconological (icon) aspects, i.e. information about the subjects, concepts and meanings of artworks, are extremely valuable for the state-of-the-art of computational tools, e.g. content recognition through computer vision. Nevertheless, a data quality evaluation for art domains, fundamental for data reuse, is still missing. The purpose of this study is filling this gap with an overview of art-historical data quality in current KGs with a focus on the icon aspects.Design/methodology/approachThis study’s analyses are based on established KG evaluation methodologies, adapted to the domain by addressing requirements from art historians’ theories. The authors first select several KGs according to Semantic Web principles. Then, the authors evaluate (1) their structures’ suitability to describe icon information through quantitative and qualitative assessment and (2) their content, qualitatively assessed in terms of correctness and completeness.FindingsThis study’s results reveal several issues on the current expression of icon information in KGs. The content evaluation shows that these domain-specific statements are generally correct but often not complete. The incompleteness is confirmed by the structure evaluation, which highlights the unsuitability of the KG schemas to describe icon information with the required granularity.Originality/valueThe main contribution of this work is an overview of the actual landscape of the icon information expressed in LOD. Therefore, it is valuable to cultural institutions by providing them a first domain-specific data quality evaluation. Since this study’s results suggest that the selected domain information is underrepresented in Semantic Web datasets, the authors highlight the need for the creation and fostering of such information to provide a more thorough art-historical dimension to LOD.
Read moreLinked Open Data: State-of-the-Art Mechanisms and Conceptual Framework
Today, one of the state-of-the-art technologies that have shown its importance towards data integration and analysis is the linked open data (LOD) systems or applications. LOD constitute of machine-readable resources or mechanisms that are useful in describing data properties. However, one of the issues with the existing systems or data models is the need for not just representing the derived information (data) in formats that can be easily understood by humans, but also creating systems that are able to process the information that they contain or support. Technically, the main mechanisms for developing the data or information processing systems are the aspects of aggregating or computing the metadata descriptions for the various process elements. This is due to the fact that there has been more than ever an increasing need for a more generalized and standard definition of data (or information) to create systems capable of providing understandable formats for the different data types and sources. To this effect, this chapter proposes a semantic-based linked open data framework (SBLODF) that integrates the different elements (entities) within information systems or models with semantics (metadata descriptions) to produce explicit and implicit information based on users’ search or queries. In essence, this work introduces a machine-readable and machine-understandable system that proves to be useful for encoding knowledge about different process domains, as well as provides the discovered information (knowledge) at a more conceptual level.
Read moreKnowledge Organization and Cultural Heritage in the Semantic Web – A Review of a Conference and a Special Journal Issue of JLIS
Review of the Knowledge Organization and Cultural Heritage: Perspectives of the Semantic Web conference held at the Academia Sinica Center for Digital Cultures in Taipei on June 2, 2016 and a special journal issue of academic papers related to the conference. Speakers came from mainland China and Taiwan across the Taiwan Strait, as well as the United States. Topics included methods and theories of Linked Open Data (LOD), the modeling of ontologies and knowledge bases, the historical and structural review of knowledge organization and representation, practical approaches used in exploring cultural assets, and the efforts of standards development for preservation and management of digital data. Speakers were encouraged to submit full papers to the peer-reviewed Journal of Library and Information Science (ISSN 0363-3640). These led to the publication of a whole special issue of the journal (vol. 43/no.1) in April 2017.
Read moreSystematic Review of Named Entity Linking and Knowledge Organisation Systems in Biomedical and Clinical Domains
Knowledge and information in the biomedical domain are predominantly represented in text format and available in diverse sources, such as articles, patents, electronic health records, and social media. Automated approaches for organising these resources face significant challenges when understanding natural language. Knowledge organisation systems serve as a crucial standard for enhancing information retrieval and sharing among humans and automated approaches. Named entity linking plays a crucial role in bridging the gap between text and knowledge organisation systems: it maps relevant entities described in text, such as diseases, chemicals, and genes, to unambiguous entries in target knowledge organisation systems that accurately describe their meaning. In this review, we analysed 102 articles published between 2013 and 2024 related to named entity linking, which provides an overview of the landscape of knowledge organisation systems in the biomedical and clinical domains and of the evolution of this task over the past decade. The findings highlight key limitations of existing approaches and identify promising directions for future research.
Read moreTrainX – Named Entity Linking with Active Sampling and Bi-Encoders
We demonstrate TrainX, a system for Named Entity Linking for medical experts. It combines state-of-the-art entity recognition and linking architectures, such as Flair and fine-tuned Bi-Encoders based on BERT, with an easy-to-use interface for healthcare professionals. We support medical experts in annotating training data by using active sampling strategies to forward informative samples to the annotator. We demonstrate that our model is capable of linking against large knowledge bases, such as UMLS (3.6 million entities), and supporting zero-shot cases, where the linker has never seen the entity before. Those zero-shot capabilities help to mitigate the problem of rare and expensive training data that is a common issue in the medical domain.
Read moreEnhancing cultural recommendations through social and linked open data
In this article, we describe a hybrid recommender system (RS) in the artistic and cultural heritage area, which takes into account the activities on social media performed by the target user and her friends, and takes advantage of linked open data (LOD) sources. Concretely, the proposed RS (1) extracts information from Facebook by analyzing content generated by users and their friends; (2) performs disambiguation tasks through LOD tools; (3) profiles the active user as a social graph; (4) provides her with personalized suggestions of artistic and cultural resources in the surroundings of the user’s current location. The last point is performed by integrating collaborative filtering algorithms with semantic technologies in order to leverage LOD sources such as DBpedia and Europeana. Based on the recommended points of cultural interest, the proposed system is also able to suggest to the active user itineraries among them, which meet her preferences and needs and are sensitive to her physical and social contexts as well. Experimental results on real users showed the effectiveness of the different modules of the proposed recommender.
Read more