- Front Matter
18
- 10.1016/j.esmoop.2022.100429
Area under the curve may hide poor generalisation to external datasets
- Apr 01, 2022
- ESMO Open
- A Kleppe
Area under the curve may hide poor generalisation to external datasets
Deep Learning applications have recently emerged for processing and analyzing medical image content. DICOM files are the raw data or the input data for Deep Learning (DL) models used to achieve various tasks such as segmentation, classification, and detection in medical diagnosis. However, such files cannot be used as specific devices produce them; they need multiple pre-processing actions due to the complexity of the problem definition when designing a DL model for medical image research. This paper introduces the innovative PreProcMed framework for data curation, medical image processing, and feature exploration. PreProcMed framework, a pioneering solution, has innovative features chained in an automated workflow using only Python technology, namely i). a new approach in data curation by automated de-identification and anonymization with the option of ii). Selection of the required MRI sequence for the DL model, followed by iii). automated conversion of 2D images in 3D volumes that are easy to use in segmentation models and fully iv). integrates annotation tools like ITKSnap and visualization tools like 3D Slicer. PreProcMed is a modular and flexible framework ready to be modified and adapted according to DL models used for medical imaging research.
Area under the curve may hide poor generalisation to external datasets
Area under the curve may hide poor generalisation to external datasets
Cross-institutional evaluation of deep learning and radiomics models in predicting microvascular invasion in hepatocellular carcinoma: validity, robustness, and ultrasound modality efficacy comparison
PurposeTo conduct a head-to-head comparison between deep learning (DL) and radiomics models across institutions for predicting microvascular invasion (MVI) in hepatocellular carcinoma (HCC) and to investigate the model robustness and generalizability through rigorous internal and external validation.MethodsThis retrospective study included 2304 preoperative images of 576 HCC lesions from two centers, with MVI status determined by postoperative histopathology. We developed DL and radiomics models for predicting the presence of MVI using B-mode ultrasound, contrast-enhanced ultrasound (CEUS) at the arterial, portal, and delayed phases, and a combined modality (B + CEUS). For radiomics, we constructed models with enlarged vs. original regions of interest (ROIs). A cross-validation approach was performed by training models on one center’s dataset and validating the other, and vice versa. This allowed assessment of the validity of different ultrasound modalities and the cross-center robustness of the models. The optimal model combined with alpha-fetoprotein (AFP) was also validated. The head-to-head comparison was based on the area under the receiver operating characteristic curve (AUC).ResultsThirteen DL models and 25 radiomics models using different ultrasound modalities were constructed and compared. B + CEUS was the optimal modality for both DL and radiomics models. The DL model achieved AUCs of 0.802–0.818 internally and 0.667–0.688 externally across the two centers, whereas radiomics achieved AUCs of 0.749–0.869 internally and 0.646–0.697 externally. The radiomics models showed overall improvement with enlarged ROIs (P < 0.05 for both CEUS and B + CEUS modalities). The DL models showed good cross-institutional robustness (P > 0.05 for all modalities, 1.6–2.1% differences in AUC for the optimal modality), whereas the radiomics models had relatively limited robustness across the two centers (12% drop-off in AUC for the optimal modality). Adding AFP improved the DL models (P < 0.05 externally) and well maintained the robustness, but did not benefit the radiomics model (P > 0.05).ConclusionCross-institutional validation indicated that DL demonstrated better robustness than radiomics for preoperative MVI prediction in patients with HCC, representing a promising solution to non-standardized ultrasound examination procedures.
Read moreBridging the gap between high-level quantum chemical methods and deep learning models
Supervised deep learning (DL) models are becoming ubiquitous in computational chemistry because they can efficiently learn complex input-output relationships and predict chemical properties at a cost significantly lower than methods based on quantum mechanics. The central challenge in many DL applications is the need to invest considerable computational resources in generating large ( N>1×105 ) training sets such that the resulting DL model can be generalized reliably to unseen systems. The lack of better alternatives has encouraged the use of low-cost and relatively inaccurate density-functional theory (DFT) methods to generate training data, leading to DL models that lack accuracy and reliability. In this article, we describe a robust and easily implemented approach based on property-specific atom-centered potentials (ACPs) that resolves this central challenge in DL model development. ACPs are one-electron potentials that are applied in combination with a computationally inexpensive but inaccurate quantum mechanical method (e.g. double-ζ DFT) and fitted against relatively few high-level data ( N≈1×103 – 1×104 ), possibly obtained from the literature. The resulting ACP-corrected methods retain the low cost of the double-ζ DFT approach, while generating high-level-quality data in unseen systems for the specific property for which they were designed. With this approach, we demonstrate that ACPs can be used as an intermediate method between high-level approaches and DL model development, enabling the calculation of large and accurate DL training sets for the chemical property of interest. We demonstrate the effectiveness of the proposed approach by predicting bond dissociation enthalpies, reaction barrier heights, and reaction energies with chemical accuracy at a computational cost lower than the DFT methods routinely used for DL training data set generation.
Read moreResearch progress in medical imaging based on deep learning of neural network
The development of computer hardware allows rapid accumulation of medical imaging data. Deep learning has shown great potential in medical imaging data analysis and establish a new area of machine learning. The commonly used deep learning models were firstly introduced in the paper, and then, summarized with the application of deep learning in the detection, classification, diagnosis, segmentation, identification of medical imaging. The application of deep learning in oral and maxillofacial radiology and other discipline of stomatology was proposed. At the end, the paper discussed the problems of deep learning in medical imaging research.
Read moreMultimodal ultrasound deep learning to detect fibrosis in early chronic kidney disease
We developed a multimodal ultrasound (US) deep learning (DL) fusion model to automatically classify early fibrosis in patients with chronic kidney disease (CKD). This prospective study included patients with CKD who underwent continuous gray-scale US, superb microvascular imaging, and strain elastography from May to November 2022. According to the pathological tubular atrophy and interstitial fibrosis score, patients were divided into minimal and mild groups (affected area ≤10% and 11 − 25% of the total cortical volume, respectively). The dataset was divided into training (70%) and test (30%) sets. A DL model combining the features of the three US modes was developed to predict early fibrosis in patients with CKD. We compared these findings with the area under the receiver operating characteristic curve (AUC) of the clinical model by analyzing the receiver operating characteristic curve in the test set. The AUC of single-mode DL based on gray-scale US, superb microvascular imaging, and strain elastography was 0.682, 0.745, and 0.648, respectively, while that of the multimodal US DL model was 0.86. The accuracy, specificity, and sensitivity of the multimodal US DL model were 0.779, 0.767, and 0.796, respectively, and the negative and positive predictive values were 0.842 and 0.706, respectively. The AUC of the multimodal US DL model was significantly better than that of the single-mode DL and clinical models. The DL algorithm developed using multimodal US images can effectively predict early fibrosis in patients with CKD with significantly greater accuracy than single-mode DL or clinical models.
Read moreA hybrid CNN and ensemble model for COVID-19 lung infection detection on chest CT scans.
COVID-19 is highly infectious and causes acute respiratory disease. Machine learning (ML) and deep learning (DL) models are vital in detecting disease from computerized chest tomography (CT) scans. The DL models outperformed the ML models. For COVID-19 detection from CT scan images, DL models are used as end-to-end models. Thus, the performance of the model is evaluated for the quality of the extracted feature and classification accuracy. There are four contributions included in this work. First, this research is motivated by studying the quality of the extracted feature from the DL by feeding these extracted to an ML model. In other words, we proposed comparing the end-to-end DL model performance against the approach of using DL for feature extraction and ML for the classification of COVID-19 CT scan images. Second, we proposed studying the effect of fusing extracted features from image descriptors, e.g., Scale-Invariant Feature Transform (SIFT), with extracted features from DL models. Third, we proposed a new Convolutional Neural Network (CNN) to be trained from scratch and then compared to the deep transfer learning on the same classification problem. Finally, we studied the performance gap between classic ML models against ensemble learning models. The proposed framework is evaluated using a CT dataset, where the obtained results are evaluated using five different metrics The obtained results revealed that using the proposed CNN model is better than using the well-known DL model for the purpose of feature extraction. Moreover, using a DL model for feature extraction and an ML model for the classification task achieved better results in comparison to using an end-to-end DL model for detecting COVID-19 CT scan images. Of note, the accuracy rate of the former method improved by using ensemble learning models instead of the classic ML models. The proposed method achieved the best accuracy rate of 99.39%.
Read moreComparison of Intratumoral and Peritumoral Deep Learning, Radiomics, and Fusion Models for Predicting KRAS Gene Mutations in Rectal Cancer Based on Endorectal Ultrasound Imaging.
We aimed at comparing intratumoral and peritumoral deep learning, radiomics, and fusion models in predicting KRAS mutations in rectal cancer using endorectal ultrasound imaging. This study included 304 patients with rectal cancer from Fujian Medical University Union Hospital. The patients were randomly divided into a training group (213 patients) and a test group (91 patients) at a 7:3 ratio. Radiomics and deep learning models were established using primary tumor and peritumoral images. In the optimally performing regions-of-interest, two fusion strategies, a feature-based and a decision-based model, were employed to build the fusion models. The Shapley additive explanation (SHAP) method was used to evaluate the significance of features in the optimal radiomics, deep learning, and fusion models. The performance of each model was assessed using the area under the receiver operating characteristic curve (AUC) and decision curve analysis (DCA). In the test cohort, both the radiomics and deep learning models exhibited optimal performance with a 10-pixel patch extension, yielding AUC values of 0.824 and 0.856, respectively. The feature-based DLRexpand10_FB model attained the highest AUC (0.896) across all study sets. In addition, the DLRexpand10_FB model demonstrated excellent sensitivity, specificity, and DCA. SHAP analysis underscored the deep learning feature (DL_1) as the most significant factor in the hybrid model. The feature-based fusion model DLRexpand10_FB can be employed to predict KRAS gene mutations based on pretreatment endorectal ultrasound images of rectal cancer. The integration of peritumoral regions enhanced the predictive performance of both the radiomics and deep learning models.
Read moreDeep Learning vs Traditional Breast Cancer Risk Models to Support Risk-Based Mammography Screening.
Deep learning breast cancer risk models demonstrate improved accuracy compared with traditional risk models but have not been prospectively tested. We compared the accuracy of a deep learning risk score derived from the patient's prior mammogram to traditional risk scores to prospectively identify patients with cancer in a cohort due for screening. We collected data on 119 139 bilateral screening mammograms in 57 617 consecutive patients screened at 5 facilities between September 18, 2017, and February 1, 2021. Patient demographics were retrieved from electronic medical records, cancer outcomes determined through regional tumor registry linkage, and comparisons made across risk models using Wilcoxon and Pearson χ2 2-sided tests. Deep learning, Tyrer-Cuzick, and National Cancer Institute Breast Cancer Risk Assessment Tool (NCI BCRAT) risk models were compared with respect to performance metrics and area under the receiver operating characteristic curves. Cancers detected per thousand patients screened were higher in patients at increased risk by the deep learning model (8.6, 95% confidence interval [CI] = 7.9 to 9.4) compared with Tyrer-Cuzick (4.4, 95% CI = 3.9 to 4.9) and NCI BCRAT (3.8, 95% CI = 3.3 to 4.3) models (P < .001). Area under the receiver operating characteristic curves of the deep learning model (0.68, 95% CI = 0.66 to 0.70) was higher compared with Tyrer-Cuzick (0.57, 95% CI = 0.54 to 0.60) and NCI BCRAT (0.57, 95% CI = 0.54 to 0.60) models. Simulated screening of the top 50th percentile risk by the deep learning model captured statistically significantly more patients with cancer compared with Tyrer-Cuzick and NCI BCRAT models (P < .001). A deep learning model to assess breast cancer risk can support feasible and effective risk-based screening and is superior to traditional models to identify patients destined to develop cancer in large screening cohorts.
Read moreEnhanced Sequence-to-Sequence Deep Transfer Learning for Day-Ahead Electricity Load Forecasting
Electricity load forecasting is a crucial undertaking within all the deregulated markets globally. Among the research challenges on a global scale, the investigation of deep transfer learning (DTL) in the field of electricity load forecasting represents a fundamental effort that can inform artificial intelligence applications in general. In this paper, a comprehensive study is reported regarding day-ahead electricity load forecasting. For this purpose, three sequence-to-sequence (Seq2seq) deep learning (DL) models are used, namely the multilayer perceptron (MLP), the convolutional neural network (CNN) and the ensemble learning model (ELM), which consists of the weighted combination of the outputs of MLP and CNN models. Also, the study focuses on the development of different forecasting strategies based on DTL, emphasizing the way the datasets are trained and fine-tuned for higher forecasting accuracy. In order to implement the forecasting strategies using deep learning models, load datasets from three Greek islands, Rhodes, Lesvos, and Chios, are used. The main purpose is to apply DTL for day-ahead predictions (1–24 h) for each month of the year for the Chios dataset after training and fine-tuning the models using the datasets of the three islands in various combinations. Four DTL strategies are illustrated. In the first strategy (DTL Case 1), each of the three DL models is trained using only the Lesvos dataset, while fine-tuning is performed on the dataset of Chios island, in order to create day-ahead predictions for the Chios load. In the second strategy (DTL Case 2), data from both Lesvos and Rhodes concurrently are used for the DL model training period, and fine-tuning is performed on the data from Chios. The third DTL strategy (DTL Case 3) involves the training of the DL models using the Lesvos dataset, and the testing period is performed directly on the Chios dataset without fine-tuning. The fourth strategy is a multi-task deep learning (MTDL) approach, which has been extensively studied in recent years. In MTDL, the three DL models are trained simultaneously on all three datasets and the final predictions are made on the unknown part of the dataset of Chios. The results obtained demonstrate that DTL can be applied with high efficiency for day-ahead load forecasting. Specifically, DTL Case 1 and 2 outperformed MTDL in terms of load prediction accuracy. Regarding the DL models, all three exhibit very high prediction accuracy, especially in the two cases with fine-tuning. The ELM excels compared to the single models. More specifically, for conducting day-ahead predictions, it is concluded that the MLP model presents the best monthly forecasts with MAPE values of 6.24% and 6.01% for the first two cases, the CNN model presents the best monthly forecasts with MAPE values of 5.57% and 5.60%, respectively, and the ELM model achieves the best monthly forecasts with MAPE values of 5.29% and 5.31%, respectively, indicating the very high accuracy it can achieve.
Read moreDeep learning for acute rib fracture detection in CT data: a systematic review and meta-analysis.
To review studies on deep learning (DL) models for classification, detection, and segmentation of rib fractures in CT data, to determine their risk of bias (ROB), and to analyse the performance of acute rib fracture detection models. Research articles written in English were retrieved from PubMed, Embase, and Web of Science in April 2023. A study was only included if a DL model was used to classify, detect, or segment rib fractures, and only if the model was trained with CT data from humans. For the ROB assessment, the Quality Assessment of Diagnostic Accuracy Studies tool was used. The performance of acute rib fracture detection models was meta-analysed with forest plots. A total of 27 studies were selected. About 75% of the studies have ROB by not reporting the patient selection criteria, including control patients or using 5-mm slice thickness CT scans. The sensitivity, precision, and F1-score of the subgroup of low ROB studies were 89.60% (95%CI, 86.31%-92.90%), 84.89% (95%CI, 81.59%-88.18%), and 86.66% (95%CI, 84.62%-88.71%), respectively. The ROB subgroup differences test for the F1-score led to a p-value below 0.1. ROB in studies mostly stems from an inappropriate patient and data selection. The studies with low ROB have better F1-score in acute rib fracture detection using DL models. This systematic review will be a reference to the taxonomy of the current status of rib fracture detection with DL models, and upcoming studies will benefit from our data extraction, our ROB assessment, and our meta-analysis.
Read moreDLBricks: Composable Benchmark Generation to Reduce Deep Learning Benchmarking Effort on CPUs
The past few years have seen a surge of applying Deep Learning (DL) models for a wide array of tasks such as image classification, object detection, machine translation, etc. While DL models provide an opportunity to solve otherwise intractable tasks, their adoption relies on them being optimized to meet latency and resource requirements. Benchmarking is a key step in this process but has been hampered in part due to the lack of representative and up-to-date benchmarking suites. This is exacerbated by the fast-evolving pace of DL models. This paper proposes DLBricks, a composable benchmark generation design that reduces the effort of developing, maintaining, and running DL benchmarks on CPUs. DLBricks decomposes DL models into a set of unique runnable networks and constructs the original model's performance using the performance of the generated benchmarks. DLBricks leverages two key observations: DL layers are the performance building blocks of DL models and layers are extensively repeated within and across DL models. Since benchmarks are generated automatically and the benchmarking time is minimized, DLBricks can keep up-to-date with the latest proposed models, relieving the pressure of selecting representative DL models. Moreover, DLBricks allows users to represent proprietary models within benchmark suites. We evaluate DLBricks using $50$ MXNet models spanning $5$ DL tasks on $4$ representative CPU systems. We show that DLBricks provides an accurate performance estimate for the DL models and reduces the benchmarking time across systems (e.g. within $95\%$ accuracy and up to $4.4\times$ benchmarking time speedup on Amazon EC2 c5.xlarge).
Read moreProspects of Deep Learning and Edge Intelligence in Agriculture
Agriculture is one of the high labor occupations around the globe. To meet the population growth and its demand, with the increase in labor cost, there is a need to explore efficient autonomous systems which may replace the traditional methods. Computer vision, edge, and deep learning (DL) models have become a promising area of research. This new paradigm of deep edge intelligence is most appropriate for agriculture activities where real-time decision-making is very important. In this chapter, the authors conduct a systematic literature review on deep learning-aided edge intelligence (EI) applications in agriculture to gather the evidence for prospects of DL at edge in agriculture. They discuss how DL models have shown outstanding performance within limited time and computation resources, and also provide future research directions to enhance the viability and applicability of complex deep learning (DL) models deployed at edge devices in agricultural applications.
Read moreExplainable deep learning diagnostic system for prediction of lung disease from medical images
Explainable deep learning diagnostic system for prediction of lung disease from medical images
Deep learning models as learners for EEG-based functional brain networks**This work was performed when Yuxuan Yang was on academic leave at Delft University of Technology.
Objective.Functional brain network (FBN) methods are commonly integrated with deep learning (DL) models for EEG analysis. Typically, an FBN is constructed to extract features from EEG data, which are then fed into a DL model for further analysis. Beyond this two-step approach, there is potential to embed FBN construction directly within DL models as a feature extraction module, enabling the models to learn EEG representations end-to-end while incorporating insights from FBNs. However, a critical prerequisite is whether DL models can effectively learn the FBN construction process.Approach.To address this, we propose using DL models to learn FBN matrices derived from EEG data. The ability of DL models to accurately reproduce these matrices would validate their capacity to learn the FBN construction process. This approach is tested on two publicly available EEG datasets, utilizing seven DL models to learn four representative FBN matrices. Model performance is assessed through mean squared error (MSE), Pearson correlation coefficient (Corr), and concordance correlation coefficient (CCC) between predicted and actual matrices.Main results.The results show that DL models achieve low MSE and relatively high Corr and CCC values when learning the Coherence network. Visualizations of predicted and error matrices reveal that while DL models capture the general structure of all four FBNs, certain regions remain difficult to model accurately. Additionally, a pairedt-test comparing global efficiency and nodal degree between predicted and actual networks indicates that most predicted networks significantly differ from the actual networks (p<0.05).Significance.These findings suggest that while DL models can learn the connectivity relationships of certain FBNs, they struggle to capture the intrinsic topological structures. This highlights the irreplaceability of traditional FBN methods in EEG analysis and underscores the need for hybrid strategies that combine FBN methods with DL models for a more comprehensive analysis.
Read moreDiversified Curriculum Innovation in College Vocal Music Education under Deep Learning Modeling
This paper takes the problems and ways of diversified curriculum construction of vocal music education in colleges and universities as the entry point and constructs a deep blended learning model of vocal music diversified curriculum based on the deep learning model. The deep evolutionary knowledge tracking model is constructed by combining the self-attention mechanism embedding model with vocal music knowledge evolution based on Transformer. To verify the effectiveness of the deep blended learning model constructed in this paper, model comparison experiments and empirical analysis were conducted. The results show that compared with the SAKT model, the DBL model in this paper has a 5.19% improvement in the average ROC-AUC value. The vocal posttest scores of the students in the experimental class were 5.09 points higher than those of the control class, with a two-tailed significance value of 0.027, which is less than 0.05 and a significant difference. This indicates that the deep learning model can effectively promote the innovative design of diversified curricula for vocal music education in colleges and universities and enhance students’ vocal music learning ability.
Read more