BEnchmarking Large Language Models for Ophthalmology (BELO): An Expert-Curated Data Set and Evaluation Framework for Knowledge and Reasoning
PurposeCurrent benchmarks evaluating large language models (LLMs) in ophthalmology are narrow and disproportionately prioritize accuracy. We introduce BEnchmarking LLMs for Ophthalmology (BELO), a standardized evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BEnchmarking LLMs for Ophthalmology assesses ophthalmology-related knowledge and reasoning quality.SubjectsThis study did not involve human participation.DesignCross-sectional study.MethodsUsing keyword matching and a fine-tuned PubMed Bidirectional Encoder Representations from Transformers model, we curated ophthalmology-specific multiple-choice questions (MCQs) from diverse medical data sets (Basic and Clinical Science Course [BCSC], Multi-Subject Multi-Choice Dataset for Medical domain [MedMCQA], Medical Question Answering [MedQA], Biomedical Semantic Indexing and Question Answering [BioASQ], and PubMed Question Answering [PubMedQA]). The data set underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by 3 senior ophthalmologists. To illustrate BELO's utility, we evaluated 8 LLMs (OpenAI o1, o3-mini, GPT-5, GPT-4o, DeepSeek-R1, MedGemma-4B, Llama-3-8B, and Gemini 1.5 Pro).Main Outcome MeasuresThe 8 LLMs were evaluated in terms using accuracy, macro-F1, and 5 text-generation metrics (Recall-Oriented Understudy for Gisting Evaluation, BERTScore, BARTScore, Metric for Evaluation of Translation with Explicit Ordering, and AlignScore). In a further evaluation involving human experts, 2 ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness.ResultsBEnchmarking LLMs for Ophthalmology consists of 900 high-quality, expert-reviewed questions aggregated from 5 sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). To demonstrate BELO's utility, we conducted a series of benchmarking exercises. In the quantitative evaluation, GPT-5 achieved the highest accuracy (0.90, 95% confidence interval [CI]: 0.89–0.92) and macro-F1 score (0.91, 95% CI: 0.89–0.93). On the other hand, the models' performance on text-generation metrics varied and were generally suboptimal, with scores ranging from 20.4 to 72.0 (out of 100, excluding the BARTScore metric), indicating room for improvement in clinical reasoning. In expert evaluations, GPT-4o was rated highest for accuracy and readability, while Gemini 1.5 Pro scored highest for completeness. A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO data set will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models.ConclusionsBEnchmarking LLMs for Ophthalmology provides a robust clinically relevant benchmark for evaluating both the accuracy and reasoning capabilities of current and emerging LLMs in ophthalmology. Future BELO benchmarking efforts will expand to include vision-based question answering and clinical scenario management tasks.Financial Disclosure(s)The authors have no proprietary or commercial interest in any materials discussed in this article.
Read more