- Supplementary Content
- 10.26434/chemrxiv.10001884/v1
OrgAIcat: Leveraging Literature Data for Enantioselectivity Prediction and Optimization in Organocatalysis
- Feb 03, 2026
- ChemRxiv
- Qi Yang + 5 more +5
The vast repository of published chemical literature constitutes a rich yet deeply flawed resource, where systemic publication bias and data heterogeneity have fostered widespread skepticism regarding its utility for predictive modeling. Here, we demonstrate that this imperfect literature data can be transformed into a robust and generalizable predictive model for asymmetric catalysis. We introduce OrgAlcat, a machine learning framework anchored by iSynth, a rigorously curated dataset of over 22,000 aminocatalytic reactions that bridge the gap between structural repositories and raw patent data. Powered by R-SPOC, a chemically interpretable descriptor that obviates computationally prohibitive quantum mechanical calculations, OrgAlcat predicts enantioselectivity with high fidelity (R 2 > 0.7, MAE < 0.30 kcal/mol) and, crucially, distinguishes intrinsic structure-selectivity relationships from mere literature trends. Validating its practical utility in a "human-in-the-loop" workflow, OrgAlcat guided the Bayesian optimization of a challenging Aldol reaction, boosting enantiomeric excess (ee) from 41% to 96% in just 12 experiments. This work establishes a validated methodology for converting historical literature into actionable intelligence, offering a powerful tool to accelerate catalyst discovery and reaction optimization.
Read more