• Home
  • Search
  • Improving Machine Learning Classification Predictions through SHAP and Features Analysis Interpretation.
  • Cite Icon2
  • https://doi.org/10.1021/acs.jcim.5c02015Copy DOI Icon

Improving Machine Learning Classification Predictions through SHAP and Features Analysis Interpretation.

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Tree-based machine learning (ML) algorithms, such as Extra Trees (ET), Random Forest (RF), Gradient Boosting Machine (GBM), and XGBoost (XGB) are among the most widely used in early drug discovery, given their versatility and performance. However, models based on these algorithms often suffer from misclassification and reduced interpretability issues, which limit their applicability in practice. To address these challenges, several approaches have been proposed, including the use of SHapley Additive Explanations (SHAP). While SHAP values are commonly used to elucidate the importance of features driving models' predictions, they can also be employed in strategies to improve their prediction performance. Building on these premises, we propose a novel approach that integrates SHAP and features value analyses to reduce misclassification in model predictions. Specifically, we benchmarked classifiers based on ET, RF, GBM, and XGB algorithms using data sets of compounds with known antiproliferative activity against three prostate cancer (PC) cell lines (i.e., PC3, LNCaP, and DU-145). The best-performing models, based on RDKit and ECFP4 descriptors with GBM and XGB algorithms, achieved MCC values above 0.58 and F1-score above 0.8 across all data sets, demonstrating satisfactory accuracy and precision. Analyses of SHAP values revealed that many misclassified compounds possess feature values that fall within the range typically associated with the opposite class. Based on these findings, we developed a misclassification-detection framework using four filtering rules, which we termed "RAW", SHAP, "RAW OR SHAP", and "RAW AND SHAP". These filtering rules successfully identified several potentially misclassified predictions, with the "RAW OR SHAP" rule retrieving up to 21%, 23%, and 63% of misclassified compounds in the PC3, DU-145, and LNCaP test sets, respectively. The developed flagging rules enable the systematic exclusion of likely misclassified compounds, even across progressively higher prediction confidence levels, thus providing a valuable approach to improve classifier performance in virtual screening applications.

Similar Papers
  • Research Article
  • Citations11

Prediction of gully erosion susceptibility through the lens of the SHapley Additive exPlanations (SHAP) method using a stacking ensemble model.

  • May 01, 2025
  • Journal of environmental management
  • Jeongho Han +2
  • Research Article

Forecasting Hematological Activity in Antiphospholipid Syndrome Using Predictive Models

  • Nov 05, 2024
  • Blood
  • Amaya Llorente +3
  • Research Article

SHAP-Enhanced Machine Learning for Explainable Stroke Risk Prediction in Hypertensive Patients

  • Dec 31, 2025
  • Journal of Material Sciences and Engineering Technology
  • David Olanrewaju Akinwale +2
  • Research Article

Factors affecting and identification of key environmental determinants of the Oncomelania hupensis snail density in the Yangtze River Delta based on machine learning models

  • Mar 04, 2026
  • Zhongguo xue xi chong bing fang zhi za zhi = Chinese journal of schistosomiasis control
  • Y Li +6
  • Research Article

Machine Learning-Based Classification of Prostate Cancer Using Clinical Biomarkers

  • Jul 29, 2025
  • International Scientific Journal of Engineering and Management
  • Dr N Veerasekar +3
  • Research Article

Artificial intelligence-based analysis of behavior and brain images in cocaine-self-administered marmosets

  • Sep 19, 2024
  • Journal of Neuroscience Methods
  • Wonmi Gu +7
  • PDF
  • Research Article
  • Citations14

AI for Automating Data Center Operations: Model Explainability in the Data Centre Context Using Shapley Additive Explanations (SHAP)

  • Apr 24, 2024
  • Electronics
  • Yibrah Gebreyesus +4
  • Preprint Article
  • Citations1

Leveraging Machine Learning for accurate and interpretable suspended sediment concentration predictions

  • May 15, 2025
  • Houda Lamane +4
  • Research Article
  • Citations5

Machine learning analysis of survival outcomes in breast cancer patients treated with chemotherapy, hormone therapy, surgery, and radiotherapy

  • Jul 10, 2025
  • Scientific Reports
  • Eyachew Misganew Tegaw +1
  • Research Article

Identifying Nuclear Data Correlated Through Predicting Bias in Integral Experiments via Applying Principal Component Analysis to Random Forest

  • Apr 01, 2025
  • Statistical Analysis and Data Mining: An ASA Data Science Journal
  • Brian Bell +3
  • PDF
  • Research Article
  • Citations7

Prediction models for postoperative recurrence of non-lactating mastitis based on machine learning

  • Apr 22, 2024
  • BMC Medical Informatics and Decision Making
  • Jiaye Sun +7
  • Research Article
  • Citations13

Prediction of plasma trough concentration of voriconazole in adult patients using machine learning

  • Jun 24, 2023
  • European Journal of Pharmaceutical Sciences
  • Lin Cheng +7
  • Research Article
  • Citations7

Exploring cement Production's role in GDP using explainable AI and sustainability analysis in Nepal

  • Jun 01, 2025
  • Case Studies in Chemical and Environmental Engineering
  • Ramhari Poudyal +7
  • Research Article

Machine Learning-Based Prediction of Institutional Delivery Dropout (IDD) Among Nigerian Women: An Exploratory Study Using SHAP Interpretability.

  • Feb 23, 2026
  • Journal of epidemiology and global health
  • Jamilu Sani +2
  • PDF
  • Research Article
  • Citations6

Machine-learning based approach to examine ecological processes influencing the diversity of riverine dissolved organic matter composition

  • May 01, 2024
  • Frontiers in Water
  • Moritz Müller +10
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.