• Home
  • Search
  • Streamlining Ophthalmic Documentation With Anonymized, Fine-Tuned Language Models: Feasibility Study.
  • https://doi.org/10.2196/72894Copy DOI Icon

Streamlining Ophthalmic Documentation With Anonymized, Fine-Tuned Language Models: Feasibility Study.

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

The growing administrative burden on clinicians, particularly in medical documentation, contributes to burnout and may compromise patient safety. Recent advancements in generative artificial intelligence (AI) offer a promising solution to improve documentation processes and address these challenges. This study aims to evaluate the feasibility of using a fine-tuned OpenAI Curie model to automate the generation of medical report summaries (epicrises) in ophthalmology. By assessing the model's performance through human and automated evaluations, this study seeks to determine its potential for reducing clinician workload while ensuring accuracy, usefulness, and compliance with regulatory requirements. A data set of around 60,000 anonymized medical letters was created using a custom algorithm to comply with General Data Protection Regulation guidelines. The Curie model was fine-tuned on this data set to generate epicrises from medical histories, diagnoses, and findings. The performance evaluation involved various human assessments and automated evaluations from 2 large language models (LLMs). In the clinical context, 49.9% (384/769) of epicrises were evaluated as helpful or excellent, whereas only 25% (194/769) were considered disturbing. In a human (manual) evaluation, formal correctness was rated significantly higher than the neutral midpoint of 2.5 on the 4-point rating scale, as determined by a 1-sample Wilcoxon signed-rank test (mean 3.59, SD 0.85; W=1686; P<.001). Using paired t tests, we found a significant reduction in time, as correcting an AI epicrisis was faster than manually writing one (mean 109.52, SD 53.30 vs mean 54.25, SD 63.34 s; t68=3.39; P<.01). While medical accuracy and usefulness showed positive trends, these did not reach statistical significance when compared to the neutral midpoint (for medical accuracy, W= 7456; P=.08), for usefulness, W=7652.5; P=.18). Epicrises generated or corrected with AI were significantly shorter than manually written ones (mean 330.43, SD 115.42 vs mean 501.07, SD 243.50 characters; t68=-6.10; P<.001). Automated LLM assessments showed alignment with human ratings, with over 52% (356/679) and 66% (489/743) of responses in the top agreement categories, respectively. This supports overall consistency, though the comparison remains a proof of concept given methodological limitations. Our study demonstrates the technical and practical feasibility of introducing fine-tuned commercial LLMs into clinical practice. The AI-generated epicrises were formally and clinically correct in many cases and showed time-saving potential. While medical accuracy and usefulness varied across cases and should be focused on in further developments, a significant workload reduction is likely. Our anonymization process showed that regulatory challenges in the context of AI with patient data can effectively be dealt with. In summary, this study highlights the promise of transformer-based LLMs in reducing administrative tasks in health care. It outlines a pipeline for integrating LLMs into European Union clinical practice, emphasizing the need for careful implementation to ensure efficiency and patient safety.

Similar Papers
  • Front Matter
  • Citations5

Overview of South Korean Guidelines for Approval of Large Language or Multimodal Models as Medical Devices: Key Features and Areas for Improvement.

  • Jan 01, 2025
  • Korean journal of radiology
  • Seong Ho Park +3
  • Research Article
  • Citations11

How do generative artificial intelligence (AI) tools and large language models (LLMs) influence language learners’ critical thinking in EFL education? A systematic review

  • Aug 04, 2025
  • Smart Learning Environments
  • Jing Liu +2
  • Front Matter
  • Citations1

Editorial: Large language models in work and business.

  • Nov 29, 2024
  • Frontiers in artificial intelligence
  • Şadi Evren Şeker
  • Research Article
  • Citations19

In-House Knowledge Management Using a Large Language Model: Focusing on Technical Specification Documents Review

  • Mar 02, 2024
  • Applied Sciences
  • Jooyeup Lee +2
  • Research Article
  • Citations7

Legal aspects of generative artificial intelligence and large language models in examinations and theses.

  • Jan 01, 2024
  • GMS journal for medical education
  • Maren März +3
  • Research Article
  • Citations1

GPT-4o and OpenAI o1 Performance on the 2024 Spanish Competitive Medical Specialty Access Examination: Cross-Sectional Quantitative Evaluation Study

  • Jan 12, 2026
  • JMIR Medical Education
  • Pau Benito +9
  • Research Article
  • Citations1

AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights

  • Oct 15, 2025
  • Proceedings of the AAAI/ACM Conference on AI Ethics and Society
  • Jiannan Xu +2
  • Research Article

Large language model use in oral and maxillofacial surgery training: a national resident survey.

  • Feb 21, 2026
  • Oral and maxillofacial surgery
  • Nolan Kranc +7
  • Research Article
  • Citations7

Global Trends in the Use of Artificial Intelligence in Dental Education: A Bibliometric Analysis.

  • May 29, 2025
  • European journal of dental education : official journal of the Association for Dental Education in Europe
  • Margarita Iniesta +1
  • Research Article
  • Citations1

Teaching Clinical Reasoning in Health Care Professions Learners Using AI-Generated Script Concordance Tests: Mixed Methods Formative Evaluation

  • Nov 20, 2025
  • JMIR Formative Research
  • Alexandre Hudon +3
  • Research Article

Evaluating gpt-4 for zero-shot classification of bleeding and clotting events: Can large language models serve as second reviewers?

  • Nov 03, 2025
  • Blood
  • Samantha Rizzo +5
  • Research Article
  • Citations16

A Review of the Opportunities and Challenges with Large Language Models in Radiology: The Road Ahead.

  • Nov 21, 2024
  • AJNR. American journal of neuroradiology
  • Neetu Soni +4
  • Research Article

Generative AI in Medical Pharmacology: Balancing Educational Benefits and Hallucination Risks

  • Apr 19, 2025
  • International Journal of Science and Research (IJSR)
  • Swati Rai +2
  • Research Article
  • Citations178

The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts

  • Jun 15, 2023
  • Research Ethics
  • Mohammad Hosseini +2
  • Research Article
  • Citations1

Bridging the Gap: Practical Challenges and Strategic Imperatives in Adopting Gen AI

  • Dec 04, 2024
  • International Conference on AI Research
  • Andrea Di Vetta
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.