Performance of artificial intelligence answering questions on disease management from people with bronchiectasis: results from the AIR-BE study.
People with bronchiectasis often seek answers to patients on the internet, including Artificial Intelligence (AI). The AIR-BE study evaluated AI-generated responses to bronchiectasis management questions formulated by European Lung Foundation (ELF) patient representatives. Three AI models - Chat-GPT, Google Bard, and Microsoft Copilot - were tested. Experts from ERS Assemblies, the ERS CONNECT CRC, and ELF patients assessed answers for accuracy (experts), comprehensiveness (experts), and understandability (patients) using a 0-10 scale. Reliability was assessed by investigators through multiple submissions of the same question. Fifteen questions of varying complexity were submitted to each AI. Responses were evaluated by 28 bronchiectasis experts and 33 patient representatives. Reliable answers were 14 for Chat-GPT, 10 for Bard, 13 for Copilot. Median accuracy scores ranged between 7.0-9.0 for Chat-GPT, 6.0-8.0 for Bard, and 6.0-8.0 for Copilot, with Chat-GPT scoring highest (8, IQR [7-9] vs 7, IQR [6-8] for both Bard and Copilot, P < 0.0001). Median comprehensiveness scores ranged from 7.5-9.0 for Chat-GPT, 6.5-9.0 for Bard, and 6.0-8.0 for Copilot, with Chat-GPT scoring highest overall (8 [7-9] vs 8 [6-9] vs 7 [6-8], P < 0.0001). Median understandability scores ranged from 8.0-9.0 for Chat-GPT, 8.0-10.0 for Bard, and 7.0-9.0 for Copilot, with both Chat-GPT and Bard (median 9, IQR [8-10] for both) outperforming Copilot (median 9, IQR [7-10]) (P = 0.0002). Despite not being a substitute for medical consultation, AI, and particularly Chat-GPT, delivers accurate, comprehensive, and understandable responses to bronchiectasis-related questions.
Read more