• Home
  • Search
  • Understanding the effect of knowledge graph extraction error on downstream graph analyses: a case study on affiliation graphs
  • https://doi.org/10.1007/s41109-025-00749-0Copy DOI Icon

Understanding the effect of knowledge graph extraction error on downstream graph analyses: a case study on affiliation graphs

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Abstract Knowledge graphs (KGs) are useful for analyzing social structures, community dynamics, institutional memberships, and other complex relationships across domains from sociology to public health. While recent advances in large language models (LLMs) have improved the scalability and accessibility of automated KG extraction from large text corpora, the impacts of extraction errors on downstream analyses are poorly understood, especially for applied scientists who depend on accurate KGs for real-world insights. To address this gap, we conducted the first evaluation of KG extraction performance at two levels: (1) micro-level edge accuracy, which is consistent with standard NLP evaluations, and manual identification of common error sources; (2) macro-level graph metrics that assess structural properties such as community detection and connectivity, which are relevant to real-world applications. Focusing on affiliation graphs of person membership in organizations extracted from social register books, our study identifies a range of extraction performance where biases across most downstream graph analysis metrics are near zero. However, as extraction performance declines, we find that many metrics exhibit increasingly pronounced biases, with each metric tending toward a consistent direction of either over- or under-estimation. Through simulations, we further show that error models commonly used in the literature do not capture these bias patterns, indicating the need for more realistic error models for KG extraction. Our findings provide actionable insights for practitioners and underscore the importance of advancing extraction methods and error modeling to ensure reliable and meaningful downstream analyses.

Similar Papers
  • Book Chapter
  • Citations2

Enhancing Manufacturing Knowledge Access with LLMs and Context-Aware Prompting

  • Oct 16, 2024
  • Frontiers in artificial intelligence and applications
  • Sebastian Monka +6
  • Research Article

Influence of Knowledge Extraction Methods on the Effectiveness of Graph-Based Rag Systems

  • Nov 26, 2025
  • NaUKMA Research Papers. Computer Science
  • Maksym Androshchuk
  • Front Matter
  • Citations1

Editorial: Large language models in work and business.

  • Nov 29, 2024
  • Frontiers in artificial intelligence
  • Şadi Evren Şeker
  • Conference Article

Quantifying Informativeness in Knowledge Graph-Augmented In-Context Learning for Multiple Choice Query Answering

  • Nov 13, 2025
  • Maryam Ghanbari +1
  • Research Article
  • Citations26

Knowledge graph–based thought: a knowledge graph–enhanced LLM framework for pan-cancer question answering

  • Jan 06, 2025
  • GigaScience
  • Yichun Feng +5
  • Research Article
  • Citations21

KNowNEt:Guided Health Information Seeking from LLMs via Knowledge Graph Integration.

  • Jan 01, 2025
  • IEEE transactions on visualization and computer graphics
  • Youfu Yan +4
  • Preprint Article
  • Citations1

Utilizing Large Language Models for Geoscience Literature Information Extraction

  • Jan 20, 2025
  • Peng Yu +3
  • Research Article
  • Citations3

ADKGD: Anomaly Detection in Knowledge Graphs with Dual-Channel Training

  • Oct 24, 2025
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Jiayang Wu +3
  • Research Article
  • Citations1

Drug repurposing for Alzheimer’s disease using a graph-of-thoughts based large language model to infer drug-disease relationships in a comprehensive knowledge graph

  • Aug 05, 2025
  • BioData Mining
  • Zhiping Paul Wang +10
  • Preprint Article

Multi-hop Reasoning and Retrieval in Embedding Space: Leveraging Large Language Models with Knowledge

  • Feb 25, 2026
  • arXiv (Cornell University)
  • Lihui Liu
  • Research Article
  • Citations1

GraphRAG-Enabled Local Large Language Model for Gestational Diabetes Mellitus: Development of a Proof-of-Concept

  • Jan 05, 2026
  • JMIR Diabetes
  • Edmund Evangelista +4
  • Research Article

The Convergence of Federated Learning, Knowledge Graphs, and Large Language Models for Language Learning: A Scoping Review

  • Mar 09, 2026
  • Applied Sciences
  • Michael Kenteris +1
  • Research Article

Stepwise Contrastive Reasoning for Retrieval-Augmented Generation over Knowledge Graphs

  • Mar 14, 2026
  • Chenxiao Lin +3
  • Research Article

HIEA: Hierarchical Inference for Entity Alignment with Collaboration of Instruction-Tuned Large Language Models and Small Models

  • Jan 18, 2026
  • Electronics
  • Xinchen Shi +2
  • Research Article

Russian-English Dataset and Entity Alignment in Knowledge Graphs with Unmatchable Entities

  • Apr 20, 2026
  • Russian Digital Libraries Journal
  • Zinaida Vladimirovna Apanovich +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.