Abstract LB373: Structuring the unstructured: How protocol quality and precision drive clinical trial completion
Abstract Introduction: Despite digitization of clinical data, drug approval processes remain prohibitively slow and 80% of trials miss enrollment timelines. Clinical enrollment and completion are influenced by both protocol design and operational factors - yet quantitative evidence linking these factors to trial outcomes remains scarce. This study developed a novel framework leveraging large language models (LLMs) and machine learning to quantify key factors such as protocol and trial complexity to assess impact on trial completion. Using a dataset of 12,096 interventional oncology trials from ClinicalTrials.gov, we transformed unstructured trial protocols into machine-readable formats, extracting key metrics such as eligibility criteria, cohort size, enrollment windows, and adherence to clinical guidelines. Methods: Eligibility criteria were abstracted and standardized using advanced machine learning (ML) techniques, including large language models (LLMs), entity recognition (NER), and unsupervised clustering using the unstructured text of trial protocols as well as guidelines from the Food and Drug Administration (FDA) and National Clinical Trials Network (NCTN). This resulted in a comprehensive database of protocol features such as eligibility criteria, interventions, endpoints, and operational details. These features were further categorized into key clinical and study domains such as cancer diagnosis, prior therapy, comorbidities, and concomitant medications. A Random Forest classifier, trained on 60% of the dataset with rigorous cross-validation, predicted trial completion with an accuracy of 81.36% (±1.39%), precision of 79.10% (±1.63%), and F1 score of 80.87% (±1.96%). Model performance was consistent across trial phases, with Phase 2 trials demonstrating the highest accuracy (82.42% ± 2.57%). Sensitivity analyses with gradient boosting and deep learning validated these findings, though the Random Forest model's interpretability was preferred. Results: Key predictors of trial success included larger target cohort size, longer recruitment windows, and number of recruitment sites, all positively correlated with completion rates. Conversely, higher protocol complexity—measured by the total number of eligibility criteria, inclusion/exclusion counts, and deviations from clinical research guidelines—was negatively associated with successful trial completion. Simplified protocols with fewer restrictive criteria and alignment with recommended guidelines significantly improved outcomes. Conclusion: This study provides a scalable, flexible framework for optimizing clinical trial design. By identifying and quantifying critical protocol features, this approach enhances trial efficiency, promotes alignment with best practices, and offers actionable insights to improve enrollment rates and patient outcomes. The methodology is extensible, enabling prospective applications to refine future trial protocols across diverse therapeutic areas. Citation Format: Ognjen Nikolic, Pranav Singh, Anne-Marie Meyer, Shannon Silkensen, Carla Rodriguez-Watson, Richard L. Schilsky. Structuring the unstructured: How protocol quality and precision drive clinical trial completion [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_2):Abstract nr LB373.
Read more