• Home
  • Search
  • Large Language Model Symptom Identification From Clinical Text: Multicenter Study
  • Cite Icon3
  • https://doi.org/10.2196/72984Copy DOI Icon

Large Language Model Symptom Identification From Clinical Text: Multicenter Study

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

BackgroundRecognizing patient symptoms is fundamental to medicine, research, and public health. However, symptoms are often underreported in coded formats even though they are routinely documented in physician notes. Large language models (LLMs), noted for their generalizability, could help bridge this gap by mimicking the role of human expert chart reviewers for symptom identification.ObjectiveThe primary objective of this multisite study was to measure the accurate identification of infectious respiratory disease symptoms using LLMs instructed to follow chart review guidelines. The secondary objective was to evaluate LLM generalizability in multisite settings without the need for site-specific training, fine-tuning, or customization.MethodsFour LLMs were evaluated: GPT-4, GPT-3.5, Llama2 70B, and Mixtral 8×7B. LLM prompts were instructed to take on the role of chart reviewers and follow symptom annotation guidelines when assessing physician notes. Ground truth labels for each note were annotated by subject matter experts. Optimal LLM prompting strategies were selected using a development corpus of 103 notes from the emergency department at Boston Children’s Hospital. The performance of each LLM was measured using a test corpus with 202 notes from Boston Children’s Hospital. The performance of an International Classification of Diseases, Tenth Revision (ICD-10)–based method was also measured as a baseline. Generalizability of the most performant LLM was then measured in a validation corpus of 308 notes from 21 emergency departments in the Indiana Health Information Exchange.ResultsSymptom identification accuracy was superior for every LLM tested for each infectious disease symptom compared to an ICD-10–based method (F1-score=45.1%). GPT-4 was the highest scoring (F1-score=91.4%; P<.001) and was significantly better than the ICD-10–based method, followed by GPT-3.5 (F1-score=90.0%; P<.001), Llama2 (F1-score=81.7%; P<.001), and Mixtral (F1-score=83.5%; P<.001). For the validation corpus, performance of the ICD-10–based method decreased (F1-score=26.9%), while GPT-4 increased (F1-score=94.0%), demonstrating better generalizability using GPT-4 (P<.001).ConclusionsLLMs significantly outperformed an ICD-10–based method for respiratory symptom identification in emergency department electronic health records. GPT-4 demonstrated the highest accuracy and generalizability, suggesting that LLMs may augment or replace traditional approaches. LLMs can be instructed to mimic human chart reviewers with high accuracy. Future work should assess broader symptom types and health care settings.

Similar Papers
  • Research Article

Evaluating gpt-4 for zero-shot classification of bleeding and clotting events: Can large language models serve as second reviewers?

  • Nov 03, 2025
  • Blood
  • Samantha Rizzo +5
  • PDF
  • Research Article
  • Citations126

Use of a Large Language Model to Assess Clinical Acuity of Adults in the Emergency Department

  • May 07, 2024
  • JAMA Network Open
  • Christopher Y K Williams +6
  • Research Article

Comparison of proprietary and fine-tuned large language models for multi-label classification of billing codes from radiology reports.

  • Mar 14, 2026
  • European radiology
  • Kamyar Arzideh +12
  • Research Article
  • Citations1

Aiding data retrieval in clinical trials with large language models: The APOLLO 11 Consortium in advanced lung cancer patients.

  • Jun 01, 2025
  • Journal of Clinical Oncology
  • Federica Corso +19
  • Preprint Article

ChatGPT Use Among Pediatric Health Care Providers: Cross-Sectional Survey Study (Preprint)

  • Jan 28, 2024
  • Susannah Kisvarday +11
  • PDF
  • Research Article
  • Citations14

ChatGPT Use Among Pediatric Health Care Providers: Cross-Sectional Survey Study

  • Sep 12, 2024
  • JMIR Formative Research
  • Susannah Kisvarday +11
  • Research Article
  • Citations2

Large Language Models Using Clinical Text in Pediatrics

  • Mar 02, 2026
  • JAMA Network Open
  • Tracy Huang +3
  • Research Article

Evaluating large language models for clinical note processing: local fine-tuning and internal-external validation using electronic health records from South Asia.

  • Feb 25, 2026
  • BMC medical informatics and decision making
  • Seyed Alireza Hasheminasab +18
  • Research Article
  • Citations2

A large language model based pipeline for extracting information from patient complaint and anamnesis in clinical notes for severity assessment

  • Jul 14, 2025
  • Scientific Reports
  • Hui Gao +12
  • Conference Article
  • Citations2

Retrieving Operation Insights with GenAI LLM: Comparative Analysis and Workflow Enhancement

  • Nov 04, 2024
  • M X Lee +1
  • Research Article
  • Citations27

Economics and Equity of Large Language Models: Health Care Perspective

  • Nov 14, 2024
  • Journal of Medical Internet Research
  • Radha Nagarajan +24
  • Research Article
  • Citations1

Abstract B006: Using large language models for scalable extraction of real-world progression events across multiple cancer types

  • Jul 10, 2025
  • Clinical Cancer Research
  • Aaron B Cohen +9
  • Conference Article

P15 Artificial Intelligence (AI) in cardiovascular rehabilitation: a scoping review of methods for assessing AI-generated healthcare content

  • Nov 01, 2025
  • Natalie Elliott +3
  • Research Article
  • Citations19

A comparative study of recent large language models on generating hospital discharge summaries for lung cancer patients.

  • Aug 01, 2025
  • Journal of biomedical informatics
  • Yiming Li +7
  • Research Article

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization.

  • Jan 01, 2025
  • Journal of registry management
  • Abhishek Shivanna +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.