- Research Article
248
- 10.1016/j.neucom.2020.07.053
Survey on Deep Neural Networks in Speech and Vision Systems
- Jul 26, 2020
- Neurocomputing
- M Alam + 4 more +4
Survey on Deep Neural Networks in Speech and Vision Systems
Novel deep neural network based pattern field classification architectures
Survey on Deep Neural Networks in Speech and Vision Systems
Survey on Deep Neural Networks in Speech and Vision Systems
On random matrices arising in deep neural networks: General I.I.D. case
We study the eigenvalue distribution of random matrices pertinent to the analysis of deep neural networks. The matrices resemble the product of the sample covariance matrices, however, an important difference is that the analog of the population covariance matrix is now a function of random data matrices (synaptic weight matrices in the deep neural network terminology). The problem has been treated in recent work [J. Pennington, S. Schoenholz and S. Ganguli, The emergence of spectral universality in deep networks, Proc. Mach. Learn. Res. 84 (2018) 1924–1932, arXiv:1802.09979] by using the techniques of free probability theory. Since, however, free probability theory deals with population covariance matrices which are independent of the data matrices, its applicability in this case has to be justified. The justification has been given in [L. Pastur, On random matrices arising in deep neural networks: Gaussian case, Pure Appl. Funct. Anal. (2020), in press, arXiv:2001.06188] for Gaussian data matrices with independent entries, a standard analytical model of free probability, by using a version of the techniques of random matrix theory. In this paper, we use another version of the techniques to extend the results of [L. Pastur, On random matrices arising in deep neural networks: Gaussian case, Pure Appl. Funct. Anal. (2020), in press, arXiv:2001.06188] to the case where the entries of the data matrices are just independent identically distributed random variables with zero mean and finite fourth moment. This, in particular, justifies the mean field approximation in the infinite width limit for the deep untrained neural networks and the property of the macroscopic universality of random matrix theory in this case.
Read moreExtracting and inserting knowledge into stacked denoising auto-encoders
Extracting and inserting knowledge into stacked denoising auto-encoders
Handwritten Digit Recognition Using Very Deep Convolutional Neural Network
Automated image classification is an essential task of the computer vision field. The tagging of images into a set of predefined groups is referred to as image classification. The implementation of computer vision to automate image classification would be beneficial because manual image evaluation and identification can be time-consuming, particularly when there are many images of different classes. Deep learning approaches are proven to overperform existing machine learning techniques in many fields in recent years, and computer vision is one of the most notable examples. The very deep neural network (VDCNN) is a powerful deep learning model for image classification, and this paper examines it briefly using MNIST handwritten digit dataset. This dataset is used to prove the efficacy of very deep neural networks over other deep learning models. The proposed study aims to comprehend the very deep neural network architecture used to accomplish a handwritten digit recognition task. The feasibility of the proposed model is evaluated using mean accuracy, validation accuracy, and standard deviation. The study results of the very deep neural network model are compared to a convolutional neural network and convolutional neural network with batch normalization. According to the results of the comparison study, very deep neural networks achieve high accuracy of 99.1% for a handwritten dataset. The outcome of the proposed work is used to interpret how well a very deep neural network performs when compared to the other two models of deep neural networks. This proposed architecture may be used to automate the classification of handwritten digits dataset.
Read moreDomain-invariant representation learning using an unsupervised domain adversarial adaptation deep neural network
Domain-invariant representation learning using an unsupervised domain adversarial adaptation deep neural network
Identifying defects and varieties of Malting Barley Kernels
This study introduces a comprehensive approach for classifying individual malting barley kernels, involving dual-sided kernel imaging, a specifically designed image processing algorithm, an optimized deep neural network architecture, and a mechanical sorting system. The proposed method achieves precise classification into multiple classes, aligning with quality standards for malting material assessment. Throughout the study, various image analysis techniques were assessed, including traditional feature engineering, established transfer learning deep neural network architectures, and our custom-designed convolutional neural network tailored for barley kernel image analysis. Comparative analysis underscores the superior performance of our network model. The study reveals that our proposed deep learning network achieves a 94% accuracy in classifying barley kernel defects and varieties, outperforming well-established transfer learning models to complex architectures that attain 93% accuracy. Additionally, it surpasses the traditional machine learning approach involving feature extraction and support vector machine classifiers, which achieve accuracy below 90% in detecting defective kernels and below 70% in varietal classification. However, we also noted the traditional approach’s advantage in morphological feature recognition. This observation guides new research toward integrating morphological feature extraction techniques with modern convolutional networks. This paper presents a deep neural network designed specifically for the analysis of cereal kernel images in two applications: defect and variety classification. It emphasizes the importance of standardizing kernel orientation and merging images from both sides of the kernel, and introduces a device for image acquisition that fulfills this need.
Read moreUnderstanding how visual information is represented in humans and machines
In the human brain, the incoming light to the retina is transformed into meaningful representations that allow us to interact with the world. In a similar vein, the RGB pixel values are transformed by a deep neural network (DNN) into meaningful representations relevant to solving a computer vision task it was trained for. Therefore, in my research, I aim to reveal insights into the visual representations in the human visual cortex and DNNs solving vision tasks. In the previous decade, DNNs have emerged as the state-of-the-art models for predicting neural responses in the human and monkey visual cortex. Research has shown that training on a task related to a brain region’s function leads to better predictivity than a randomly initialized network. Based on this observation, we proposed that we can use DNNs trained on different computer vision tasks to identify functional mapping of the human visual cortex. To validate our proposed idea, we first investigate a brain region occipital place area (OPA) using DNNs trained on scene parsing task and scene classification task. From the previous investigations about OPA’s functions, we knew that it encodes navigational affordances that require spatial information about the scene. Therefore, we hypothesized that OPA’s representation should be closer to a scene parsing model than a scene classification model as the scene parsing task explicitly requires spatial information about the scene. Our results showed that scene parsing models had representation closer to OPA than scene classification models thus validating our approach. We then selected multiple DNNs performing a wide range of computer vision tasks ranging from low-level tasks such as edge detection, 3D tasks such as surface normals, and semantic tasks such as semantic segmentation. We compared the representations of these DNNs with all the regions in the visual cortex, thus revealing the functional representations of different regions of the visual cortex. Our results highly converged with previous investigations of these brain regions validating the feasibility of the proposed approach in finding functional representations of the human brain. Our results also provided new insights into underinvestigated brain regions that can serve as starting hypotheses and promote further investigation into those brain regions. We applied the same approach to find representational insights about the DNNs. A DNN usually consists of multiple layers with each layer performing a computation leading to the final layer that performs prediction for a given task. Training on different tasks could lead to very different representations. Therefore, we first investigate at which stage does the representation in DNNs trained on different tasks starts to differ. We further investigate if the DNNs trained on similar tasks lead to similar representations and on dissimilar tasks lead to more dissimilar representations. We selected the same set of DNNs used in the previous work that were trained on the Taskonomy dataset on a diverse range of 2D, 3D and semantic tasks. Then, given a DNN trained on a particular task, we compared the representation of multiple layers to corresponding layers in other DNNs. From this analysis, we aimed to reveal where in the network architecture task-specific representation is prominent. We found that task specificity increases as we go deeper into the DNN architecture and similar tasks start to cluster in groups. We found that the grouping we found using representational similarity was highly correlated with grouping based on transfer learning thus creating an interesting application of the approach to model selection in transfer learning. During previous works, several new measures were introduced to compare DNN representations. So, we identified the commonalities in different measures and unified different measures into a single framework referred to as duality diagram similarity. This work opens up new possibilities for similarity measures to understand DNN representations. While demonstrating a much higher correlation with transfer learning than previous state-of-the-art measures we extend it to understanding layer-wise representations of models trained on the Imagenet and Places dataset using different tasks and demonstrate its applicability to layer selection for transfer learning. In all the previous works, we used the task-specific DNN representations to understand the representations in the human visual cortex and other DNNs. We were able to interpret our findings in terms of computer vision tasks such as edge detection, semantic segmentation, depth estimation, etc. however we were not able to map the representations to human interpretable concepts. Therefore in our most recent work, we developed a new method that associates individual artificial neurons with human interpretable concepts. Overall, the works in this thesis revealed new insights into the representation of the visual cortex and DNNs...
Read moreDetecting Hate Speech in Tweets Using Different Deep Neural Network Architectures
One of the major problems, apparent in online social media, is the toxic online content. This has continued unabated, as people from diverse cultural backgrounds access the Internet, concealing their identity under the cloud of anonymity. Deep neural networks have been employed to detect hate speech from online content. This paper describes three different Deep Neural Network (DNN) Architectures for detection of hate words in Twitter - Gated Recurrent Unit (GRU), useful in capturing sequence orders, Convolution Neural Network (CNN), good for feature extraction, and Universal Language Model Fine-tuning (ULMFiT) model, which is based on transfer learning technique. ULMFiT model uses the DNN Architecture called Average-SGD Weight-Dropped Long Short Term Memory (AWD-LSTM). AWD –LSTM model was pre-trained using WikiText103 dataset. This method significantly outperformed the other Architectures.
Read moreStochasticNet: Forming Deep Neural Networks via Stochastic Connectivity
Deep neural networks are a branch in machine learning that has seen a meteoric rise in popularity due to its powerful abilities to represent and model high-level abstractions in highly complex data. One area in deep neural networks that are ripe for exploration is neural connectivity formation. A pivotal study on the brain tissue of rats found that synaptic formation for specific functional connectivity in neocortical neural microcircuits can be surprisingly well modeled and predicted as a random formation. Motivated by this intriguing finding, we introduce the concept of StochasticNet where deep neural networks are formed via stochastic connectivity between neurons. As a result, any type of deep neural networks can be formed as a StochasticNet by allowing the neuron connectivity to be stochastic. Stochastic synaptic formations in a deep neural network architecture can allow for efficient utilization of neurons for performing specific tasks. To evaluate the feasibility of such a deep neural network architecture, we train a StochasticNet using four different image datasets (CIFAR-10, MNIST, SVHN, and STL-10). Experimental results show that a StochasticNet using less than half the number of neural connections as a conventional deep neural network achieves comparable accuracy and reduces overfitting on the CIFAR-10, MNIST, and SVHN data sets. Interestingly, StochasticNet with less than half the number of neural connections, achieved a higher accuracy (relative improvement in test error rate of $\sim 6$ % compared to ConvNet) on the STL-10 data set than a conventional deep neural network. Finally, the StochasticNets have faster operational speeds while achieving better or similar accuracy performances.
Read moreNNReArch: A Tensor Program Scheduling Framework Against Neural Network Architecture Reverse Engineering
Architecture reverse engineering has become an emerging attack against deep neural network (DNN) implementations. Several prior works have utilized side-channel leakage to recover the model architecture while the an DNN is executing on a hardware acceleration platform. In this work, we target an open-source deep-learning accelerator, Versatile Tensor Accelerator (VTA), and utilize electromagnetic (EM) side-channel leakage to comprehensively learn the association between DNN architecture configurations and EM emanations. We also consider the holistic system–including the low-level tensor program code of the VTA accelerator on a Xilinx FPGA, and explore the effect of such low-level configurations on the EM leakage. Our study demonstrates that both the optimization and configuration of tensor programs will affect the EM side-channel leakage.Gaining knowledge of the association between low-level tensor program and the EM emanations, we propose NNReArch, a lightweight tensor program scheduling framework against side-channel-based DNN model architecture reverse engineering. Specifically, NNReArch targets reshaping the EM traces of different DNN operators, through scheduling the tensor program execution of the DNN model so as to confuse the adversary. NNReArch is a comprehensive protection framework supporting two modes, a balanced mode that strikes a balance between the DNN model confidentiality and execution performance, and a secure mode where the most secure setting is chosen. We implement and evaluate the proposed framework on the open-source VTA with state-of-the-art DNN architectures. The experimental results demonstrate that NNReArch can efficiently enhance the model architecture security with a small performance overhead. In addition, the proposed obfuscation technique makes reverse engineering of the DNN architecture significantly harder.
Read moreThe Use of Transfer Learning with Very Deep Convolutional Neural Network in Quality Management
Purpose: The aim of the article is to develop an algorithm for classifying cracks in the analyzed images using modern methods of deep machine learning and transfer learning based on pretrained convolutional neural network - Inception-ResNet-v2. Design/Methodology/Approach: Transfer learning based on the pretrained convolutional neural network was used to categorize the images. The fully conected layer of the Inception-ResNet-v2 network has been modified. The last layer was trained using a two-class (binary) linear SVM (Support Vector Machine). In total, 20,000 training cases (images) were used to train the fully connected layer within transfer learning process. The research analyzed the possibility of using the deep neural networks for quick and fully automatic identification of cracks / defects on the surface of analyzed parts. Findings: The results indicate that pretrained convolutional neural network using SVM to train a fully connected layer is a very effective solution for visual crack / fault detection. In the analyzed model, a positive classification was obtained at the level of 99.89%. Practical Implications: The model presented in the article can be used in quality control carried out by process monitoring. An effective model for identifying defective parts can be used in both logistics and production processes. Originality/Value: A novelty is the use of a freely available, deep neural network trained to classify 1000 categories of various images for binary categorization of faults (cracks). The algorithm was adjusted by replacing the primary, 1000-output fully connected layer in the Inception-ResNet-v2 network with a binary layer (2 categories). The fully connected layer has been trained using the classification version of the popular SVM learner, but thanks to the combination of this layer with the sophisticated fearure extraction ability of the pre-trained Inception-ResNet-v2 deep network, the resulting predictive model enables the classification of defects with a very high level of accuracy.
Read moreComparison of Deep Learning and Traditional Machine Learning Models for Predicting Mild Cognitive Impairment Using Plasma Proteomic Biomarkers.
Mild cognitive impairment (MCI) is a clinical condition characterized by a decline in cognitive ability and progression of cognitive impairment. It is often considered a transitional stage between normal aging and Alzheimer's disease (AD). This study aimed to compare deep learning (DL) and traditional machine learning (ML) methods in predicting MCI using plasma proteomic biomarkers. A total of 239 adults were selected from the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort along with a pool of 146 plasma proteomic biomarkers. We evaluated seven traditional ML models (support vector machines (SVMs), logistic regression (LR), naïve Bayes (NB), random forest (RF), k-nearest neighbor (KNN), gradient boosting machine (GBM), and extreme gradient boosting (XGBoost)) and six variations of a deep neural network (DNN) model-the DL model in the H2O package. Least Absolute Shrinkage and Selection Operator (LASSO) selected 35 proteomic biomarkers from the pool. Based on grid search, the DNN model with an activation function of "Rectifier With Dropout" with 2 layers and 32 of 35 selected proteomic biomarkers revealed the best model with the highest accuracy of 0.995 and an F1 Score of 0.996, while among seven traditional ML methods, XGBoost was the best with an accuracy of 0.986 and an F1 Score of 0.985. Several biomarkers were correlated with the APOE-ε4 genotype, polygenic hazard score (PHS), and three clinical cerebrospinal fluid biomarkers (Aβ42, tTau, and pTau). Bioinformatics analysis using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) revealed several molecular functions and pathways associated with the selected biomarkers, including cytokine-cytokine receptor interaction, cholesterol metabolism, and regulation of lipid localization. The results showed that the DL model may represent a promising tool in the prediction of MCI. These plasma proteomic biomarkers may help with early diagnosis, prognostic risk stratification, and early treatment interventions for individuals at risk for MCI.
Read moreA Short Review of Deep Learning Neural Networks in Protein Structure Prediction Problems
Determining the structure of a protein given its sequence is a challenging problem. Deep learning is a rapidly evolving field which excels at problems where there are complex relationships between input features and desired outputs. Deep Neural Networks have become popular for solving problems in protein science. Various deep neural network architectures have been proposed including deep feed-forward neural networks, recurrent neural networks and more recently neural Turing machines and memory networks. This article provides a short review of deep learning applied to protein prediction problems.
Read moreA Cognitive Level Evaluation Method Based on a Deep Neural Network for Online Learning: From a Bloom’s Taxonomy of Cognition Objectives Perspective
The evaluation of the learning process is an effective way to realize personalized online learning. Real-time evaluation of learners’ cognitive level during online learning helps to monitor learners’ cognitive state and adjust learning strategies to improve the quality of online learning. However, most of the existing cognitive level evaluation methods use manual coding or traditional machine learning methods, which are time-consuming and laborious. They cannot fully mine the implicit cognitive semantic information in unstructured text data, making the cognitive level evaluation inefficient. Therefore, this study proposed the bidirectional gated recurrent convolutional neural network combined with an attention mechanism (AM-BiGRU-CNN) deep neural network cognitive level evaluation method, and based on Bloom’s taxonomy of cognition objectives, taking the unstructured interactive text data released by 9167 learners in the massive open online course (MOOC) forum as an empirical study to support the method. The study found that the AM-BiGRU-CNN method has the best evaluation effect, with the overall accuracy of the evaluation of the six cognitive levels reaching 84.21%, of which the F1-Score at the creating level is 91.77%. The experimental results show that the deep neural network method can effectively identify the cognitive features implicit in the text and can be better applied to the automatic evaluation of the cognitive level of online learners. This study provides a technical reference for the evaluation of the cognitive level of the students in the online learning environment, and automatic evaluation in the realization of personalized learning strategies, teaching intervention, and resources recommended have higher application value.
Read moreDiversified radar micro-Doppler simulations as training data for deep residual neural networks
A key challenge in radar micro-Doppler classification is the difficulty in obtaining a large amount of training data due to costs in time and human resources. Small training datasets limit the depth of deep neural networks (DNNs), and, hence, attainable classification accuracy. In this work, a novel method for diversifying Kinect-based motion capture (MOCAP) simulations of human micro-Doppler to span a wider range of potential observations, e.g. speed, body size, and style, is proposed. By applying three transformations, a small set of MOCAP measurements is expanded to generate a large training dataset for network initialization of a 30-layer deep residual neural network. Results show that the proposed training methodology and residual DNN yield improved bottleneck feature performance and the highest overall classification accuracy among other DNN architectures, including transfer learning from the 1.5 million sample ImageNet database.
Read more