Abstract 1107: Feature selection and machine learning strategies optimize an affordable molecular assay for cholangiocarcinoma subtype
Background: Cholangiocarcinoma is a molecularly heterogenous cancer arising from the biliary epithelium. Individual omic techniques have identified clinically relevant, targetable subtypes but most tumors have no targetable alterations based on exome or transcriptomic evaluation. Integrated multiomics approaches incorporating transcriptomic, proteomic, and phosphoproteomic characterization provide deeper understanding and identify non-mutated, activated pathways that can be targeted. To optimize the potential clinical utility of these approaches, minimal classifier features must be defined. Methods: Approximately 60,000 RNA, protein, and phosphoprotein features were extracted from 35 cholangiocarcinomas treated at the primary US study site. 206 patients treated and profiled at an institution in China were also included. Using the Multiomics Factor Analysis clustering method, patients were sorted into 3 distinct molecular profiles. Important molecular features were selected using competing feature selection (FS) methods: Boruta, RreliefF, variable selection using random forests (RF), recursive factor elimination (RFE), median decrease accuracy in RF, and Concrete Autoencoders. Machine learning classification prediction models including Extreme Gradient Boosting, gradient-boosted trees, multinomial logit, and support vector machine (SVM) were tested on each FS. Each combination of FS and predictive model was tested with 10 replicates of 10-fold cross-validation of a randomly-selected balanced test set and an oversampled, balanced training subset. Accuracy, F1 score, AUC, precision, and recall were assessed for each FS and predictive model combination. A custom function of physical assay cost and classification accuracy was optimized to find the best assay. KEGG pathway overrepresentation analysis was used to analyze feature subsets. Results: A consistent set of 50 important RNA and protein features was identified, including previously reported prognostic factors, like SpryD4 protein expression, as well as previously unreported pathway surrogates, such as Slc2A2 or Pipox protein. Pathway overrepresentation analysis found insulin signaling and resistance and amino acid biosynthesis/metabolism pathways were overrepresented in the discriminative feature set. The gradient-boosted trees algorithm combined with the RFE FS method had the best performance, with a multiclass macro average (±sd) cross-validation accuracy 89% (±6%), F1 score 89% (±8%), AUC 0.98 (±0.02), precision 90% (±10%), recall 89% (±6%). Conclusions: Feature selection methods and predictive machine learning models work synchronously to develop accurate molecular subtype assays based on just 50 protein and RNA levels. This strategy has the potential to make molecular subtyping for cholangiocarcinoma and other molecularly complex cancers fast and widely available. Citation Format: Ellen Larson, Erik Jessen, Dong-Gi Mun, Jennifer Tomlinson, Amro Abdelrahman, Danielle Carlson, Hojjat Salehinejad, Rory Smoot. Feature selection and machine learning strategies optimize an affordable molecular assay for cholangiocarcinoma subtype [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 1107.
Read more