- Conference Article
- 10.1109/ccwc67433.2026.11393803
Machine Learning Approach for Cancer Detection, Subtyping, and Progression Using TP53 Gene Mutation Signatures
- Jan 05, 2026
- Anoushka Vijay
The integration of artificial intelligence (AI) is transforming medical research and the healthcare system with its role in the management of complex and global medical issues such as cancer. AI can improve operational efficiency to reduce the burden on the medical professional. Integration of AI into the flow from early diagnosis to treatment and remission can be beneficial to efficiently manage various diseases and help improve patient life expectancy. Integration of AI, Machine Learning (ML) and Deep Learning (DL) in healthcare to build accurate, efficient, and effective tools to solve complex healthcare challenges can open the path to making healthcare accessible, affordable, and equitable throughout the world. Gene expressions and regulations are the key to cell development, DNA damage repair, genomic integrity, and responses to drug therapies. The TP53 gene, which is also known as the 'Guardian of Genomes', regulates the production of the p53 protein. Various environmental and bodily stresses can trigger the synthesis of mutated p53, a common finding in many types of cancers. Many studies focus on molecular understanding of mutations, genomic networks, and mutation effects. We propose a machine learning (ML) framework to build generic predictive models for classification, sub-typing, and progression of cancers using mutation-based TP53 data. The primary goal of our study is to showcase the potential of using TP53 based data in building ML tools that can be useful for cancer research and diagnosis. One key challenge to achieve this goal is to find a comprehensive TP53 data set which includes important features to train the models and provide useful results. However, most publicly available datasets are curated from previous publications and lack uniformity or correlation. In this work, we use the curated TP53 mutation data set to build and train ML frameworks and models after applying data processing techniques to infer cancer information. Through data augmentation and feature selection, we achieved <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$>80 \%$</tex> accuracy in both classification and subtyping. The progression model to predict tumor grade and cancer stage achieves an accuracy of around <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$> 4 8 \%$</tex>. Our study identifies gaps in the available data that can be addressed to improve the model results in the future. The data augmentation can result in data leakage, thereby producing overly accurate results, the limitation that are discussed in the paper. Our approach underscores the potential of integrating multiple types of TP53 genomic data with machine learning models to build useful analytical tools for cancer research.
Read more