• Home
  • Search
  • Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search
  • Open Access IconOpen Access
  • Cite Icon5
  • https://doi.org/10.1109/bigdata62323.2024.10825160Copy DOI Icon

Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search

  • Dec 15, 2024
  • Edward Kim +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Creation and curation of knowledge graphs at scale can be used to exponentially accelerate the discovery, matching, and analysis of diseases in real-world data. While disease ontologies are useful for annotation, integration, and analysis of biological data, codified disease and procedure categories e.g. SNOMED-CT, ICD10, CPT, etc. rarely capture all of the nuances in a patient condition or, in the case of rare disease, may not even exist. Furthermore, there are multiple disease definitions used in data sources and publications, each having its own structure and hierarchy. Mapping between ontologies, finding disease clusters, and building a representation of the chosen disease area are resource-intensive, often requiring significant human capital. We propose the creation and curation of a patient knowledge graph utilizing large language model extraction techniques. In order to expand in volume and scale, knowledge graphs with generalized language capability allow for data to be extracted using natural language rather than being constrained by the exact terminology or hierarchy of existing ontologies. We develop a method of mapping back to existing ontologies such as MeSH, SNOMED-CT, RxNORM, HPO, etc. to ground the extracted entities to known entities in the medical community. We have access to one of the largest ambulatory care EHR databases in the country. To demonstrate the effectiveness of our method, we benchmark our extraction in a test set with over 33.6M unique patients, in the area of patient search. In this case study, we perform a patient search for a rare disease: Dravet syndrome. Dravet syndrome was codified as an ICD10 recognizable disease in October 2020. In the following research, we describe our method of the construction of patient-specific knowledge graphs and subsequent searches for patients who exhibit symptoms of a particular disease. Using patients with confirmed ICD10 codes for Dravet syndrome as our ground truth, we utilize our LLM-based entity extraction techniques and formalize an algorithmic way of characterizing patients in a grounded ontology to assist in mapping patients to specific diseases. Finally, we present the results of a real-world discovery method on Beta-propeller protein-associated neurodegeneration (BPAN), identifying patients with a rare disease, where no ground truth currently exists.

Similar Papers
  • Research Article
  • Citations4

A web server for automatic analysis and extraction of relevant biological knowledge

  • May 25, 2007
  • Computers in Biology and Medicine
  • Juan Cedano +6
  • Conference Article

DsOn: Ontology-Driven Model for Symptom and Drug Knowledge Extraction on Social Media

  • Jan 01, 2020
  • Farahnazgolrooy Motlagh
  • Research Article
  • Citations3

A Representation-Based Methodology for Developing High-Value Knowledge Engineering Software Products: Theory, Application, and Implementation

  • Sep 12, 2013
  • Journal of Computing and Information Science in Engineering
  • S Desa +1
  • Research Article
  • Citations13

Nursing students' experiences with the objective structured clinical examination (OSCE): A qualitative study

  • Jan 01, 2022
  • International Journal of Africa Nursing Sciences
  • Yosra Raziani +2
  • PDF
  • Research Article

Knowledge Extraction for Sleep Apnea Medical Diagnosis

  • Jan 01, 2016
  • Journal of Health Education Research & Development
  • Hung Hsiang Chiu +1
  • Research Article
  • Citations10

Genome data classification based on fuzzy matching

  • Oct 13, 2012
  • CSI Transactions on ICT
  • Nagamma Patil +2
  • Conference Article
  • Citations24

Intelligent sampling for big data using bootstrap sampling and chebyshev inequality

  • May 01, 2014
  • Ashwin Satyanarayana
  • Book Chapter
  • Citations2

Benchmarking Internet of Things Deployment: Frameworks, Best Practices, and Experiences

  • Jan 01, 2015
  • Franck Le Gall +5
  • Conference Article
  • Citations13

A Ubiquitous Supervisory System Based on Social Context Awareness

  • Jan 01, 2008
  • Takuo Suganuma +4
  • Research Article
  • Citations1

Legal counselling system

  • Feb 01, 1994
  • Sadhana
  • M Shashi +2
  • Research Article
  • Citations2

Critical Multiliteration Model Based on Project Based Learning Approach in Developing Basic School of Metacognition Thinking Skills

  • Dec 11, 2019
  • International Journal of Science and Applied Science: Conference Series
  • Ani Hendriyani +3
  • Book Chapter
  • Citations1

Entity Correspondence with Second-Order Markov Logic

  • Jan 01, 2013
  • Ying Xu +5
  • PDF
  • Research Article
  • Citations2

Construction and effectiveness evaluation of a knowledge-based infectious disease monitoring and decision support system

  • Aug 14, 2023
  • Scientific Reports
  • Mengying Wang +5
  • Conference Instance
  • Citations7

Proceedings of the 11th Knowledge Capture Conference

  • Dec 02, 2021
  • Research Article
  • Citations26

FCANN: A new approach for extraction and representation of knowledge from ANN trained via Formal Concept Analysis

  • May 09, 2008
  • Neurocomputing
  • Luis E Zárate +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.