- Front Matter
18
- 10.1016/j.esmoop.2022.100429
Area under the curve may hide poor generalisation to external datasets
- Apr 01, 2022
- ESMO Open
- A Kleppe
Area under the curve may hide poor generalisation to external datasets
Applying deep learning models to RNA-Seq data poses substantial challenges, primarily due to the high dimensionality of the data and the limited sample sizes. To address these issues, this study introduces an advanced deep learning pipeline that integrates feature engineering with data augmentation. The engineering application focuses on biomedical engineering, specifically the classification of RNA-Seq datasets for disease diagnosis. The proposed framework was initially validated on synthetic datasets generated from Naive Bayes, where MLP-based augmentation yielded a notable improvement in predictive performance. Building on this foundation, we applied the approach to chromophobe renal cell carcinoma (KICH) RNA-Seq data from The Cancer Genome Atlas (TCGA). Following standard preprocessing steps normalization, transformation, and dimensionality reduction, the analysis concentrated on three main aspects: augmentation strategies, preprocessing methods, and explainable AI (XAI) techniques in relation to classification outcomes. Feature selection was performed through PCA, Boruta, and RF-based methods. Three augmentation strategies linear interpolation, SMOTE, and MixUp were evaluated. To maintain methodological rigor, augmentation was applied exclusively to the training set, while the test set was held out for unbiased evaluation. Within this framework, we conducted a comparative assessment of multiple deep learning architectures, including MLP, GNN, and the recently proposed Kolmogorov-Arnold networks (KAN). The GNN achieved the highest classification accuracy (99.47%) when trained with MixUp augmentation combined with RF feature selection, and achieved the best F1 score (0.9948). Consequently, the GNN-based XAI framework was applied to the RF dataset enriched with MixUp. XAI analyses identified the top 20 most influential genes, such as HNF4A, DACH2, MAPK15, and NAT2, which played the greatest role in classification, thereby confirming the biological plausibility of the model outputs. To further validate model robustness, cervical cancer and Alzheimer's RNA-Seq datasets were also tested, yielding consistent and reliable results. Overall, the findings highlight the value of incorporating data augmentation into deep learning models for RNA-Seq analysis, not only to improve predictive performance but also to enhance biological interpretability through explainable AI approaches.
Area under the curve may hide poor generalisation to external datasets
Area under the curve may hide poor generalisation to external datasets
Abstract 1817: Differential expression of long non-coding RNA in colon adenocarcinoma RNA-sequence data set
Introduction: Colon cancer is the fourth most common cancer in the United States and the third leading cause of cancer-related death. High throughput genomic sequencing has led to a number of significant advances in tumor biology and in identifying novel signaling molecules such as long non-coding RNA (lncRNA). The aim of this study was to identify differentially expressed lncRNA in colon cancer from a large RNA-sequencing (RNA-seq) data set. Methods: The raw RNA-seq files of 398 patients with colon cancer were downloaded from The Cancer Genome Atlas (TCGA). Sequencing files were aligned using STAR (Spliced Transcripts Alignment to a Reference) to the most recent genome annotation. Using a subset of patients with paired colon cancer and normal colon epithelium RNA-seq data (n=40), an exploratory binomial regression model was used to calculate differential RNA expression. The most differentially expressed lncRNA were identified from the exploratory analysis and verified by comparing the larger colon cancer RNA-seq data set (n=358) with the normal colon epithelium RNA-seq data set (n=40). Results: 33,514 genes were identified from the comparative analysis, and using differential expression cut off values of > +1.5 or < -1.5 log fold change and a false discovery rate of <0.05; 543 were upregulated and 1822 were downregulated. Within this dysregulated group, 60 lncRNAs were identified. Using the larger data set (n=358 vs n=40), 41/60 lncRNA remained differentially expressed, of which 15 were downregulated and 26 were upregulated. Twenty-four of these lncRNAs have not previously been described in colon cancer, 12 of which are upregulated and 12 of which are downregulated (Table 1). Conclusions: This analysis of RNA-seq data from TCGA has identified dysregulated lncRNAs which have not previously been described in human colon cancer. These lncRNAs may have significant roles in colon cancer tumor signaling. Table 1.Differential expression of lncRNA not described in colon cancerlncRNALog Fold ChangeP-valueDescription in other cancersKRT16P14.1041.36E-21Increased in lung squamous cell carcinomaBOK-AS13.7784.98E-09Increased in oral squamous cell carcinomaCELP3.1151.30E-09Not availableCLDN10-AS12.3944.61E-12Not availablePOU6F2-AS12.2678.25E-09Not availableSLC7A11-AS12.1931.60E-13Decreased in gastric adenocarcinomaCSAG42.1083.59E-10Not availableLUCAT12.0224.21E-11Increased in Ovarian cancer, renal clear cell carcinoma, head and neck squamous cell carcinoma, Non-small cell cancer, glioma, osteosarcoma, esophageal squamous cell carcinomaVPS9D1-AS11.9795.27E-21Decreased in gastric adenocarcinoma, Increased in non-small cell lung cancerLEF1-AS11.8464.48E-13Increased in glioblastomaCASC81.6485.64E-06Single nucleotide polymorphisms in colorectal adenocarcinomaLINC004911.5903.11E-04Not availableLINC00648-1.5021.02E-15Not availableLINC00163-1.6751.03E-22Not availableSATB2-AS1-1.8074.18E-09Increased in osteosarcomaLINC00702-2.1424.81E-18Not availableTRHDE-AS1-2.2623.44E-24Not availableLINC00908-2.3082.76E-103Not availableLINC00483-2.4012.87E-39Increased in gastric adenocarcinomaLINC00461-2.4401.42E-41Increased in glioma, multiple myelomaFENDRR-2.6872.12E-58Decreased in breast cancer, prostate cancer, osteosarcoma, gastric adenocarcinoma,ARHGEF26-AS1-2.7148.55E-32Not availableADAMTS9-AS2-2.9041.78E-74Increased in lung cancer salivary adenoid cystic carcinoma, Decreased in glioma, gastric adenocarcinomaMT1JP-4.4601.08E-114Decreased in retinoblastoma, Decreased in gastric adenocarcinoma Citation Format: Stephen J. O'Brien, Theodore Kalbfleisch, Sudhir Srivastava, Shesh Rai, Susan Galandiuk. Differential expression of long non-coding RNA in colon adenocarcinoma RNA-sequence data set [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 1817.
Read moreApplication of artificial intelligence model in pathological staging and prognosis of clear cell renal cell carcinoma
This study aims to develop a deep learning (DL) model based on whole-slide images (WSIs) to predict the pathological stage of clear cell renal cell carcinoma (ccRCC). The histopathological images of 513 ccRCC patients were downloaded from The Cancer Genome Atlas (TCGA) database and randomly divided into training set and validation set according to the ratio of 8∶2. The CLAM algorithm was used to establish the DL model, and the stability of the model was evaluated in the external validation set. DL features were extracted from the model to construct a prognostic risk model, which was validated in an external dataset. The results showed that the DL model showed excellent prediction ability with an area under the curve (AUC) of 0.875 and an average accuracy score of 0.809, indicating that the model could reliably distinguish ccRCC patients at different stages from histopathological images. In addition, the prognostic risk model constructed by DL characteristics showed that the overall survival rate of patients in the high-risk group was significantly lower than that in the low-risk group (P = 0.003), and AUC values for predicting 1-, 3- and 5-year overall survival rates were 0.68, 0.69 and 0.69, respectively, indicating that the prediction model had high sensitivity and specificity. The results of the validation set are consistent with the above results. Therefore, DL model can accurately predict the pathological stage and prognosis of ccRCC patients, and provide certain reference value for clinical diagnosis.
Read moreData from Multifactorial Deep Learning Reveals Pan-Cancer Genomic Tumor Clusters with Distinct Immunogenomic Landscape and Response to Immunotherapy
<div>AbstractPurpose:<p>Tumor genomic features have been of particular interest because of their potential impact on the tumor immune microenvironment and response to immunotherapy. Due to the substantial heterogeneity, an integrative approach incorporating diverse molecular features is needed to characterize immunologic features underlying primary resistance to immunotherapy and for the establishment of novel predictive biomarkers.</p>Experimental Design:<p>We developed a pan-cancer deep machine learning model integrating tumor mutation burden, microsatellite instability, and somatic copy-number alterations to classify tumors of different types into different genomic clusters, and assessed the immune microenvironment in each genomic cluster and the association of each genomic cluster with response to immunotherapy.</p>Results:<p>Our model grouped 8,646 tumors of 29 cancer types from The Cancer Genome Atlas into four genomic clusters. Analysis of RNA-sequencing data revealed distinct immune microenvironment in tumors of each genomic class. Furthermore, applying this model to tumors from two melanoma immunotherapy clinical cohorts demonstrated that patients with melanoma of different genomic classes achieved different benefit from immunotherapy. Interestingly, tumors in cluster 4 demonstrated a cold immune microenvironment and lack of benefit from immunotherapy despite high microsatellite instability burden.</p>Conclusions:<p>Our study provides a proof for principle that deep learning modeling may have the potential to discover intrinsic statistical cross-modality correlations of multifactorial input data to dissect the molecular mechanisms underlying primary resistance to immunotherapy, which likely involves multiple factors from both the tumor and host at different molecular levels.</p></div>
Read moreA hybrid CNN and ensemble model for COVID-19 lung infection detection on chest CT scans.
COVID-19 is highly infectious and causes acute respiratory disease. Machine learning (ML) and deep learning (DL) models are vital in detecting disease from computerized chest tomography (CT) scans. The DL models outperformed the ML models. For COVID-19 detection from CT scan images, DL models are used as end-to-end models. Thus, the performance of the model is evaluated for the quality of the extracted feature and classification accuracy. There are four contributions included in this work. First, this research is motivated by studying the quality of the extracted feature from the DL by feeding these extracted to an ML model. In other words, we proposed comparing the end-to-end DL model performance against the approach of using DL for feature extraction and ML for the classification of COVID-19 CT scan images. Second, we proposed studying the effect of fusing extracted features from image descriptors, e.g., Scale-Invariant Feature Transform (SIFT), with extracted features from DL models. Third, we proposed a new Convolutional Neural Network (CNN) to be trained from scratch and then compared to the deep transfer learning on the same classification problem. Finally, we studied the performance gap between classic ML models against ensemble learning models. The proposed framework is evaluated using a CT dataset, where the obtained results are evaluated using five different metrics The obtained results revealed that using the proposed CNN model is better than using the well-known DL model for the purpose of feature extraction. Moreover, using a DL model for feature extraction and an ML model for the classification task achieved better results in comparison to using an end-to-end DL model for detecting COVID-19 CT scan images. Of note, the accuracy rate of the former method improved by using ensemble learning models instead of the classic ML models. The proposed method achieved the best accuracy rate of 99.39%.
Read moreExplainable deep learning diagnostic system for prediction of lung disease from medical images
Explainable deep learning diagnostic system for prediction of lung disease from medical images
Intelligent skin disease prediction system using transfer learning and explainable artificial intelligence
Skin diseases impact millions of people around the world and pose a severe risk to public health. These diseases have a wide range of effects on the skin’s structure, functionality, and appearance. Identifying and predicting skin diseases are laborious processes that require a complete physical examination, a review of the patient’s medical history, and proper laboratory diagnostic testing. Additionally, it necessitates a significant number of histological and clinical characteristics for examination and subsequent treatment. As a disease’s complexity and quantity of features grow, identifying and predicting it becomes more challenging. This research proposes a deep learning (DL) model utilizing transfer learning (TL) to quickly identify skin diseases like chickenpox, measles, and monkeypox. A pre-trained VGG16 is used for transfer learning. The VGG16 can identify and predict diseases more quickly by learning symptom patterns. Images of the skin from the four classes of chickenpox, measles, monkeypox, and normal are included in the dataset. The dataset is separated into training and testing. The experimental results performed on the dataset demonstrate that the VGG16 model can identify and predict skin diseases with 93.29% testing accuracy. However, the VGG16 model does not explain why and how the system operates because deep learning models are black boxes. Deep learning models’ opacity stands in the way of their widespread application in the healthcare sector. In order to make this a valuable system for the health sector, this article employs layer-wise relevance propagation (LRP) to determine the relevance scores of each input. The identified symptoms provide valuable insights that could support timely diagnosis and treatment decisions for skin diseases.
Read moreCross-institutional evaluation of deep learning and radiomics models in predicting microvascular invasion in hepatocellular carcinoma: validity, robustness, and ultrasound modality efficacy comparison
PurposeTo conduct a head-to-head comparison between deep learning (DL) and radiomics models across institutions for predicting microvascular invasion (MVI) in hepatocellular carcinoma (HCC) and to investigate the model robustness and generalizability through rigorous internal and external validation.MethodsThis retrospective study included 2304 preoperative images of 576 HCC lesions from two centers, with MVI status determined by postoperative histopathology. We developed DL and radiomics models for predicting the presence of MVI using B-mode ultrasound, contrast-enhanced ultrasound (CEUS) at the arterial, portal, and delayed phases, and a combined modality (B + CEUS). For radiomics, we constructed models with enlarged vs. original regions of interest (ROIs). A cross-validation approach was performed by training models on one center’s dataset and validating the other, and vice versa. This allowed assessment of the validity of different ultrasound modalities and the cross-center robustness of the models. The optimal model combined with alpha-fetoprotein (AFP) was also validated. The head-to-head comparison was based on the area under the receiver operating characteristic curve (AUC).ResultsThirteen DL models and 25 radiomics models using different ultrasound modalities were constructed and compared. B + CEUS was the optimal modality for both DL and radiomics models. The DL model achieved AUCs of 0.802–0.818 internally and 0.667–0.688 externally across the two centers, whereas radiomics achieved AUCs of 0.749–0.869 internally and 0.646–0.697 externally. The radiomics models showed overall improvement with enlarged ROIs (P < 0.05 for both CEUS and B + CEUS modalities). The DL models showed good cross-institutional robustness (P > 0.05 for all modalities, 1.6–2.1% differences in AUC for the optimal modality), whereas the radiomics models had relatively limited robustness across the two centers (12% drop-off in AUC for the optimal modality). Adding AFP improved the DL models (P < 0.05 externally) and well maintained the robustness, but did not benefit the radiomics model (P > 0.05).ConclusionCross-institutional validation indicated that DL demonstrated better robustness than radiomics for preoperative MVI prediction in patients with HCC, representing a promising solution to non-standardized ultrasound examination procedures.
Read morePredicting craniofacial fibrous dysplasia growth status: an exploratory study of a hybrid radiomics and deep learning model based on computed tomography images
Predicting craniofacial fibrous dysplasia growth status: an exploratory study of a hybrid radiomics and deep learning model based on computed tomography images
Read moreExplainable breast cancer prediction from 3-dimensional dynamic contrast-enhanced magnetic resonance imaging
Deep learning models have been instrumental in extracting critical indicators for breast cancer diagnosis - the prevalent malignancy among women worldwide - from baseline magnetic resonance imaging. However, many existing models do not fully leverage the rich spatial information available in the 3D structure of medical imaging data, potentially overlooking important contextual details. This develops an explainable deep learning framework for classifying breast cancer that leverages the complete 3D and provides classification results alongside visual explanations of the decision-making process. The preprocessing pipeline is fed with 3D sequences containing ‘tumour’ and ‘non-tumour’ regions. It includes a 3D Adaptive Unsharp Mask (AUM) filter to reduce noise and augment image class, followed by normalisation and data augmentation. Classification is then achieved by training an augmented ResNet150 model. Three explainable artificial intelligence (XAI) techniques, including Shapley Additive Explanations, 3D Gradient-Weighted Class Activation Mapping, and Contextual Importance and Utility, are employed to provide improved interpretability. The model demonstrates state-of-the-art performance over the QIN-BREAST dataset, achieving testing accuracies of 98.861% for ‘tumours’ and 99.447% for ‘non-tumours’, as well as over the Duke Breast Cancer Dataset, where it achieves 99.104% for ‘tumours’ and 99.753% for ‘non-tumours’, while offering enhanced interpretability through XAI methods.
Read moreEnhanced Sequence-to-Sequence Deep Transfer Learning for Day-Ahead Electricity Load Forecasting
Electricity load forecasting is a crucial undertaking within all the deregulated markets globally. Among the research challenges on a global scale, the investigation of deep transfer learning (DTL) in the field of electricity load forecasting represents a fundamental effort that can inform artificial intelligence applications in general. In this paper, a comprehensive study is reported regarding day-ahead electricity load forecasting. For this purpose, three sequence-to-sequence (Seq2seq) deep learning (DL) models are used, namely the multilayer perceptron (MLP), the convolutional neural network (CNN) and the ensemble learning model (ELM), which consists of the weighted combination of the outputs of MLP and CNN models. Also, the study focuses on the development of different forecasting strategies based on DTL, emphasizing the way the datasets are trained and fine-tuned for higher forecasting accuracy. In order to implement the forecasting strategies using deep learning models, load datasets from three Greek islands, Rhodes, Lesvos, and Chios, are used. The main purpose is to apply DTL for day-ahead predictions (1–24 h) for each month of the year for the Chios dataset after training and fine-tuning the models using the datasets of the three islands in various combinations. Four DTL strategies are illustrated. In the first strategy (DTL Case 1), each of the three DL models is trained using only the Lesvos dataset, while fine-tuning is performed on the dataset of Chios island, in order to create day-ahead predictions for the Chios load. In the second strategy (DTL Case 2), data from both Lesvos and Rhodes concurrently are used for the DL model training period, and fine-tuning is performed on the data from Chios. The third DTL strategy (DTL Case 3) involves the training of the DL models using the Lesvos dataset, and the testing period is performed directly on the Chios dataset without fine-tuning. The fourth strategy is a multi-task deep learning (MTDL) approach, which has been extensively studied in recent years. In MTDL, the three DL models are trained simultaneously on all three datasets and the final predictions are made on the unknown part of the dataset of Chios. The results obtained demonstrate that DTL can be applied with high efficiency for day-ahead load forecasting. Specifically, DTL Case 1 and 2 outperformed MTDL in terms of load prediction accuracy. Regarding the DL models, all three exhibit very high prediction accuracy, especially in the two cases with fine-tuning. The ELM excels compared to the single models. More specifically, for conducting day-ahead predictions, it is concluded that the MLP model presents the best monthly forecasts with MAPE values of 6.24% and 6.01% for the first two cases, the CNN model presents the best monthly forecasts with MAPE values of 5.57% and 5.60%, respectively, and the ELM model achieves the best monthly forecasts with MAPE values of 5.29% and 5.31%, respectively, indicating the very high accuracy it can achieve.
Read moreAutomated segmentation of dental restorations using deep learning: exploring data augmentation techniques.
Automated segmentation of dental restorations using deep learning: exploring data augmentation techniques.
Enhancing brain tumor detection in MRI images through explainable AI using Grad-CAM with Resnet 50
This study addresses the critical challenge of detecting brain tumors using MRI images, a pivotal task in medical diagnostics that demands high accuracy and interpretability. While deep learning has shown remarkable success in medical image analysis, there remains a substantial need for models that are not only accurate but also interpretable to healthcare professionals. The existing methodologies, predominantly deep learning-based, often act as black boxes, providing little insight into their decision-making process. This research introduces an integrated approach using ResNet50, a deep learning model, combined with Gradient-weighted Class Activation Mapping (Grad-CAM) to offer a transparent and explainable framework for brain tumor detection. We employed a dataset of MRI images, enhanced through data augmentation, to train and validate our model. The results demonstrate a significant improvement in model performance, with a testing accuracy of 98.52% and precision-recall metrics exceeding 98%, showcasing the model’s effectiveness in distinguishing tumor presence. The application of Grad-CAM provides insightful visual explanations, illustrating the model’s focus areas in making predictions. This fusion of high accuracy and explainability holds profound implications for medical diagnostics, offering a pathway towards more reliable and interpretable brain tumor detection tools.
Read moreDesign and Implementation of a Custom Lightweight CNN with Integrated XAI for Diagnostic Skin Lesion Analysis
The differentiation of benign and malignant skin lesion types is an essential problem in medical images, which facilitates diagnosis and treatment of skin cancer. Traditional deep learning approaches usually depend on the use of the pre-trained models, which are not customized to the specific dataset or task. In this work, to tackle the challenge, a Custom Residual-based Depth-wise Separable Lightweight Inception Model with Custom Attention Mechanism is proposed together with Explainable AI (XAI) techniques. It is written from scratch, using only bottleneck convolutions (1x1) and spatially separable convolutions (1x3,3x1), which makes it really lightweight and efficient. The approach includes developing a new model architecture based on residuals and multi-scale (which known as Inception-style) feature extraction using a novel attention mechanism that guides the network on the important image features. Furthermore, XAI methods such as Grad-CAM, LIME, and SHAP are employed to analyze and explain the model's output, facilitating transparency and reliability in the medical domain. This coursework train and test the model on an accessible skin lesion dataset, using extensive data preprocessing, data augmentation, and class balancing. The results show that the proposed model is superior, with highest accuracy, precision, recall, and F1-score as well as good ROC-AUC and PR-AUC. The XAI visualizations also further confirm the model’s attention on clinically relevant regions of the images, which makes the model more interpretable. In conclusion, a deep learning model that is lightweight, interpretable, and efficient for skin lesion classification is proposed. The combination of PI layers and XAI methodologies guarantees a good level of performance and transparency, thus proving to be an effective instrument for the analysis of medical images.
Read moreMultimodal ultrasound deep learning to detect fibrosis in early chronic kidney disease
We developed a multimodal ultrasound (US) deep learning (DL) fusion model to automatically classify early fibrosis in patients with chronic kidney disease (CKD). This prospective study included patients with CKD who underwent continuous gray-scale US, superb microvascular imaging, and strain elastography from May to November 2022. According to the pathological tubular atrophy and interstitial fibrosis score, patients were divided into minimal and mild groups (affected area ≤10% and 11 − 25% of the total cortical volume, respectively). The dataset was divided into training (70%) and test (30%) sets. A DL model combining the features of the three US modes was developed to predict early fibrosis in patients with CKD. We compared these findings with the area under the receiver operating characteristic curve (AUC) of the clinical model by analyzing the receiver operating characteristic curve in the test set. The AUC of single-mode DL based on gray-scale US, superb microvascular imaging, and strain elastography was 0.682, 0.745, and 0.648, respectively, while that of the multimodal US DL model was 0.86. The accuracy, specificity, and sensitivity of the multimodal US DL model were 0.779, 0.767, and 0.796, respectively, and the negative and positive predictive values were 0.842 and 0.706, respectively. The AUC of the multimodal US DL model was significantly better than that of the single-mode DL and clinical models. The DL algorithm developed using multimodal US images can effectively predict early fibrosis in patients with CKD with significantly greater accuracy than single-mode DL or clinical models.
Read more