P0748 Using Large Language Models to Reconcile International IBD Guidelines and Generate Consensus Statements
Abstract Background Clinical guidelines for Inflammatory Bowel Disease (IBD) are essential for standardizing care, but their length, technical language, and inconsistent recommendations make them difficult for busy clinicians to use at the point of care. Manually comparing and reconciling multiple guidelines is labor-intensive and often impractical in real-world clinical practice. We aimed to develop and evaluate a proof-of-concept tool using large language models (LLMs) with retrieval-augmented generation (RAG) to improve guideline readability by harmonizing recommendations, highlighting consensus and controversy, and generating concise, actionable statements. Methods We designed an LLM-RAG pipeline (GPT-4o) to automatically segment guideline documents into manageable units, enrich them with metadata and summaries, and retrieve relevant content in response to clinical queries. The system synthesizes recommendations across guidelines, presenting areas of consensus and disagreement in a structured format. Four major international IBD guidelines (ACG, ECCO, BSG, ACPGBI) were analyzed across eight common clinical questions spanning Crohn’s disease and ulcerative colitis. Tool-generated outputs were benchmarked against expert summaries and evaluated by four independent reviewers using 5-point Likert scales for completeness, accuracy, relevance, coherence, and conciseness. Results The tool consistently improved guideline readability by distilling complex text into shorter, structured responses. It achieved mean scores of 4.34 (95% CI, 4.20–4.48) for consensus recognition and 4.61 (95% CI, 4.46–4.77) for disagreement detection. Completeness, accuracy, and relevance all scored >4.0. Although conciseness was lower (3.84, 95% CI, 3.50–4.19), reviewers noted that outputs captured essential information while substantially reducing textual burden. Outline generation performance was moderate (3.25, 95% CI, 2.85–3.65), reflecting challenges in extracting all relevant subtopics. In 7 of 8 clinical scenarios (87.5%), the tool’s recommendations aligned with expert conclusions. Conclusion This proof-of-concept study demonstrates that an LLM-RAG framework can systematically reconcile international IBD guidelines and present them in a more readable, clinically usable format. By reducing complexity and making consensus and controversy explicit, such tools can help clinicians access key evidence more efficiently, support faster decision-making at the bedside, and reduce practice variation. With further refinement, this approach could contribute to “living guidelines” that are continuously updated and more accessible to end-users, ultimately enhancing patient care.
Read more