- Research Article
- 10.1158/1538-7445.am2025-5084
Abstract 5084: Evaluation of single-cell foundation models for cancer outcome predictions
- Apr 21, 2025
- Cancer Research
- Haitham Elmarakeby + 3 more +3
Introduction: Tumor single-cell data provide unique insights into the cellular dynamics of cancer progression and therapeutic outcomes. Inspired by advancements in natural language processing (NLP), foundation models are increasingly used in single-cell analyses to uncover biological insights and identify clinical biomarkers. However, the performance of these models on cancer-specific tasks remains underexplored. We present an extensive evaluation of foundation models trained on millions of single-cell profiles for cancer-focused investigations. Methods: We evaluated nine predictive models-three traditional methods and six foundation models (four Geneformer and two scGPT models), across various downstream tasks. Traditional models used embeddings from high variable genes (HVG), principal components (PCA), and generative model embedding (scVI), each paired with a random forest classifier for downstream classification. The evaluated foundation models varied in architecture sizes, pretraining dataset sizes (30M to 95M cell profiles), and training datasets, which were either generic or cancer-specific. We assessed the quality of representations by measuring their predictive performance on lung cancer treatment status (TKI treated vs. treatment-naive patients) and breast cancer subtypes (ER+ vs. TNBC). We examined how model characteristics influenced their performance in these domains for all tasks. Results: Zero-shot evaluations of the foundation models showed superior single-cell data clustering, measured by average silhouette width (ASW), compared to traditional methods (HVG, PCA, scVI), suggesting that pre-trained embeddings from millions of single-cell profiles may offer enhanced representations. We thus evaluated the learned representation on downstream predictive tasks, showing that foundation models with a larger pretraining dataset (95M cells), cancer-specific training, larger input size (4096 genes), and deeper model (12 layers) achieved the best performance among foundation models (F1=0.97) for predicting treatment status in lung cancer. In contrast, smaller and generic foundation models had significantly lower performance, however, additional self-supervised training on task-specific data rescued their performance to match larger foundation models. Similar trends were observed in predicting breast cancer subtypes. Notably, the best foundation model’s performance was not significantly different from the best traditional method (PCA) in predicting breast cancer subtypes. Conclusions: These results highlight key factors driving foundation model performance on cancer-specific tasks. While current foundation models show inconsistent contributions to cancer-specific questions, observed performance trends highlight their potential to enhance single-cell data representation and drive significant advancements in cancer research and precision medicine. Citation Format: Haitham Elmarakeby, Ahmed Roman, Shreya Johri, Eliezer M. Van Allen. Evaluation of single-cell foundation models for cancer outcome predictions [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 5084.
Read more