- Research Article
534
- 10.1016/j.neunet.2014.08.005
Deep Convolutional Neural Networks for Large-scale Speech Tasks
- Sep 16, 2014
- Neural Networks
- Tara N Sainath + 6 more +6
Deep Convolutional Neural Networks for Large-scale Speech Tasks
Deep Learning, developed from the multi-layer ANN concept, is a machine learning method that compiles every detail of the learning process to obtain more abstract data, multi-level data, and more complex data features through the composition of various mathematical functions. Deep MLP, another name for DNN (Deep Neural Network), is MLP, which has more than three layers. DNN's ability will increase in solving problems when more layers are used. Another advantage of DNN is the variety of layer types used, including fully connected layers, convolution layers, softmax layers, recurrent layers, and others. Autoencoder is a type of ANN that trained to reconstruct input patterns in such a way that the output of the deepest hidden layer is a vector resulting from the reduction of dimensions of the input pattern. This study proposes a novel DNN architecture called Deep Auto-Encoder Semi Convolutional Neural Network (DAESCNN) to improve the performance of conventional ANN performance. Annual rainfall data will be used to test forecast performance using DAESCNN. In this study, annual rainfall data were obtained from weather stations of Samarinda city, East Kalimantan, Indonesia in the period 2006–2016. In this study, all 2006–2014 data points were used as training input data, all 2007–2015 data points were used as training target data, while 2016 data points were used as validation of training outcomes. The DAESCNN training process executed through a program built using MATLAB. The training process also carried out with varying training parameters to show its performance. The results shown that DAESCNN generally has excellent performance, is above 99%.
Deep Convolutional Neural Networks for Large-scale Speech Tasks
Deep Convolutional Neural Networks for Large-scale Speech Tasks
Evolutionary pruning of transfer learned deep convolutional neural network for breast cancer diagnosis in digital breast tomosynthesis
Deep learning models are highly parameterized, resulting in difficulty in inference and transfer learning for image recognition tasks. In this work, we propose a layered pathway evolution method to compress a deep convolutional neural network (DCNN) for classification of masses in digital breast tomosynthesis (DBT). The objective is to prune the number of tunable parameters while preserving the classification accuracy. In the first stage transfer learning, 19 632 augmented regions-of-interest (ROIs) from 2454 mass lesions on mammograms were used to train a pre-trained DCNN on ImageNet. In the second stage transfer learning, the DCNN was used as a feature extractor followed by feature selection and random forest classification. The pathway evolution was performed using genetic algorithm in an iterative approach with tournament selection driven by count-preserving crossover and mutation. The second stage was trained with 9120 DBT ROIs from 228 mass lesions using leave-one-case-out cross-validation. The DCNN was reduced by 87% in the number of neurons, 34% in the number of parameters, and 95% in the number of multiply-and-add operations required in the convolutional layers. The test AUC on 89 mass lesions from 94 independent DBT cases before and after pruning were 0.88 and 0.90, respectively, and the difference was not statistically significant (p > 0.05). The proposed DCNN compression approach can reduce the number of required operations by 95% while maintaining the classification performance. The approach can be extended to other deep neural networks and imaging tasks where transfer learning is appropriate.
Read moreMultimodal and Crossmodal Representation Learning from Textual and Visual Features with Bidirectional Deep Neural Networks for Video Hyperlinking
Video hyperlinking represents a classical example of multimodal problems. Common approaches to such problems are early fusion of the initial modalities and crossmodal translation from one modality to the other. Recently, deep neural networks, especially deep autoencoders, have proven promising both for crossmodal translation and for early fusion via multimodal embedding. A particular architecture, bidirectional symmetrical deep neural networks, have been proven to yield improved multimodal embeddings over classical autoencoders, while also being able to perform crossmodal translation. In this work, we focus firstly at evaluating good single-modal continuous representations both for textual and for visual information. Word2Vec and paragraph vectors are evaluated for representing collections of words, such as parts of automatic transcripts and multiple visual concepts, while different deep convolutional neural networks are evaluated for directly embedding visual information, avoiding the creation of visual concepts. Secondly, we evaluate methods for multimodal fusion and crossmodal translation, with different single-modal pairs, in the task of video hyperlinking. Bidirectional (symmetrical) deep neural networks were shown to successfully tackle downsides of multimodal autoencoders and yield a superior multimodal representation. In this work, we extensively tests them in different settings, with different single-modal representations, within the context of video-hyperlinking. Our novel bidirectional symmetrical deep neural networks are compared to classical autoencoders and are shown to yield significantly improved multimodal embeddings that significantly (alpha=0.0001) outperform multimodal embeddings obtained by deep autoencoders with an absolute improvement in precision at 10 of 14.1% when embedding visual concepts and automatic transcripts and an absolute improvement of 4.3% when embedding automatic transcripts with features obtained with very deep convolutional neural networks, yielding 80% of precision at 10.
Read moreHandwritten Digit Recognition Using Very Deep Convolutional Neural Network
Automated image classification is an essential task of the computer vision field. The tagging of images into a set of predefined groups is referred to as image classification. The implementation of computer vision to automate image classification would be beneficial because manual image evaluation and identification can be time-consuming, particularly when there are many images of different classes. Deep learning approaches are proven to overperform existing machine learning techniques in many fields in recent years, and computer vision is one of the most notable examples. The very deep neural network (VDCNN) is a powerful deep learning model for image classification, and this paper examines it briefly using MNIST handwritten digit dataset. This dataset is used to prove the efficacy of very deep neural networks over other deep learning models. The proposed study aims to comprehend the very deep neural network architecture used to accomplish a handwritten digit recognition task. The feasibility of the proposed model is evaluated using mean accuracy, validation accuracy, and standard deviation. The study results of the very deep neural network model are compared to a convolutional neural network and convolutional neural network with batch normalization. According to the results of the comparison study, very deep neural networks achieve high accuracy of 99.1% for a handwritten dataset. The outcome of the proposed work is used to interpret how well a very deep neural network performs when compared to the other two models of deep neural networks. This proposed architecture may be used to automate the classification of handwritten digits dataset.
Read moreMaxout neurons for deep convolutional and LSTM neural networks in speech recognition
Maxout neurons for deep convolutional and LSTM neural networks in speech recognition
DCNN-based Ship Classification using Enhanced Edge Information and Inception Module
The excellent feature extraction ability of deep convolutional neural networks (DCNNs) has been demonstrated in many image processing tasks, by which image classification can achieve high accuracy with only raw input images. However, the specific image features that influence the classification results are not readily determinable and what lies behind the predictions is unclear. This study proposes a method combining the Sobel and Canny operators and an Inception module for ship classification. The Sobel and Canny operators obtain enhanced edge features from the input images. A convolutional layer is replaced with the Inception module, which can automatically select the proper convolution kernel for ship objects in different image regions. The principle is that the high-level features abstracted by the DCNN, and the features obtained by multi-convolution concatenation of the Inception module must ultimately derive from the edge information of the preprocessing input images. This indicates that the classification results are based on the input edge features, which indirectly interpret the classification results to some extent. Experimental results show that the combination of the edge features and the Inception module improves DCNN ship classification performance. The original model with the raw dataset has an average accuracy of 88.72%, while when using enhanced edge features as input, it achieves the best performance of 90.54% among all models. The model that replaces the fifth convolutional layer with the Inception module has the best performance of 89.50%. It performs close to VGG-16 on the raw dataset and is significantly better than other deep neural networks. The results validate the functionality and feasibility of the idea posited.
Read moreAcoustic feature-based sentiment analysis of call center data
With the advancement of machine learning methods, audio sentiment analysis has become an active research area in recent years. For example, business organizations are interested in persuasion tactics from vocal cues and acoustic measures in speech. A typical approach is to find a set of acoustic features from audio data that can indicate or predict a customer's attitude, opinion, or emotion state. For audio signals, acoustic features have been widely used in many machine learning applications, such as music classification, language recognition, emotion recognition, and so on. For emotion recognition, previous work shows that pitch and speech rate features are important features. This thesis work focuses on determining sentiment from call center audio records, each containing a conversation between a sales representative and a customer. The sentiment of an audio record is considered positive if the conversation ended with an appointment being made, and is negative otherwise. In this project, a data processing and machine learning pipeline for this problem has been developed. It consists of three major steps: 1) an audio record is split into segments by speaker turns; 2) acoustic features are extracted from each segment; and 3) classification models are trained on the acoustic features to predict sentiment. Different set of features have been used and different machine learning methods, including classical machine learning algorithms and deep neural networks, have been implemented in the pipeline. In our deep neural network method, the feature vectors of audio segments are stacked in temporal order into a feature matrix, which is fed into deep convolution neural networks as input. Experimental results based on real data shows that acoustic features, such as Mel frequency cepstral coefficients, timbre and Chroma features, are good indicators for sentiment. Temporal information in an audio record can be captured by deep convolutional neural networks for improved prediction accuracy.
Read moreAbsolute distance measurement based on laser self-mixing interferometry and deep neural network
Self-mixing interferometry (SMI) is superior to other laser interferometry methods due to its simplicity and compactness. However, SMI signals are often complex and difficult to process due to interference in the form of variations in the effective reflectivity of the target, noisy signals, complex signal shapes, and other dependencies. Deep neural networks have been a very popular area of research in computer artificial intelligence in recent years, allowing more implicit features to be uncovered than traditional shallow machine learning. It has been shown that convolutional and back-propagation neural networks can be used for SMI signal processing. There are also studies that have used machine learning genetic algorithms for absolute distance measurement. Based on the above, this study used convolutional neural networks to form a deep neural network for absolute distance measurement based on SMI technology. We first trained the deep convolutional neural network at different feedback strengths. The results of the Convolutional Neural Network (CNN) model showed a coefficient of determination of 0.9987. which is consistent with the required model. The trained network was then used to estimate absolute distances with and without the addition of noise. The comparison proves that the proposed method is noise-proof and has high adaptability for measurements under different conditions.
Read moreConvolutional neural network-based wind pressure prediction on low-rise buildings
Convolutional neural network-based wind pressure prediction on low-rise buildings
Exploring Convolution Neural Network for Branch Prediction
Recently, there have been significant advances in deep neural networks (DNNs) and they have shown distinctive performance in speech recognition, natural language processing, and image recognition. In this paper, we explore DNNs to push the limit for branch prediction. We treat branch prediction as a classification problem and employ both deep convolutional neural networks (CNNs), ranging from LeNet to ResNet-50, and deep belief network (DBN) for branch prediction. We compare the effectiveness of DNNs with the state-of-the-art branch predictors, including the perceptron, our prior work, Multi-poTAGE+SC, and MTAGE+SC branch predictors. The last two are the most recent winners of championship branch prediction (CBP) contests. Several interesting observations emerged from our study. First, for branch prediction, the DNNs outperform the perceptron model as high as 60-80%. Second, we analyze the impact of the depth of CNNs (i.e., number of convolutional layers and pooling layers) on the misprediction rates. The results confirm that deeper CNN structures can lead to lower misprediction rates. Third, the DBN could outperform our prior work, but not outperform the state-of-the-art TAGE-like branch predictor; the ResNet-50 could not only outperform our prior work, but also the Multi-poTAGE+SC and MTAGE+SC.
Read moreApplying a Deep Learning Neural Network to Gait-Based Pedestrian Automatic Detection and Recognition
Gait recognition is a noncontact biometric procedure that determines the identity or health status of a person by analyzing his or her walking posture and habits, including skeletal and joint movements. The most remarkable feature of this method is the possibility of conducting recognition without demanding much cooperation from participants. Therefore, this recognition technique has attracted much attention from scholars. Additionally, because of the rapid development of graphics processing unit technology, related hardware and computation performance, the applications of deep-learning technology are considerably enhanced. The objective of this study was to apply a deep neural network (DNN), which employs deep-learning technology, to achieve gait-based automatic pedestrian detection and recognition. In contrast to using wearable devices to precisely capture skeletal and joint movements, pedestrian color-image sequences were used as input in this study. Subsequently, a pretraining convolutional neural network (CNN) was employed to capture pedestrian location and extract pedestrian dense optical flow to serve as concrete low-level feature inputs. Then, a finely-tuned DNN based on the wide residual network was employed to extract high-level abstract features. In addition, to overcome the difficulty of obtaining local temporal features by using a 2D CNN, part of the 3D convolutional structure was introduced into the CNN. This design enabled use of limited memory to acquire more effective features and enhance the DNN performance. The experimental results show that the proposed method has exceptional performance for pedestrian detection and recognition.
Read moreAdversarial Robustness of Deep Convolutional Neural Network-based Image Recognition Models: A Review
Deep convolutional neural networks have achieved great success in recent years. They have been widely used in various applications such as optical and SAR image scene classification, object detection and recognition, semantic segmentation, and change detection. However, deep neural networks rely on large-scale high-quality training data, and can only guarantee good performance when the training and test data are independently sampled from the same distribution. Deep convolutional neural networks are found to be vulnerable to subtle adversarial perturbations. This adversarial vulnerability prevents the deployment of deep neural networks in security-sensitive applications such as medical, surveillance, autonomous driving and military scenarios. This paper first presents a holistic view of security issues for deep convolutional neural network-based image recognition systems. The entire information processing chain is analyzed regarding safety and security risks. In particular, poisoning attacks and evasion attacks on deep convolutional neural networks are analyzed in detail. The root causes of adversarial vulnerabilities of deep recognition models are also discussed. Then, we give a formal definition of adversarial robustness and present a comprehensive review of adversarial attacks, adversarial defense, and adversarial robustness evaluation. Rather than listing existing research, we focus on the threat models for the adversarial attack and defense arms race. We perform a detailed analysis of several representative adversarial attacks on SAR image recognition models and provide an example of adversarial robustness evaluation. Finally, several open questions are discussed regarding recent research progress from our workgroup. This paper can be further used as a reference to develop more robust deep neural network-based image recognition models in dynamic adversarial scenarios.
Read moreComputer-Assisted Differential Diagnosis of Pyoderma Gangrenosum and Venous Ulcers with Deep Neural Networks.
(1) Background: Pyoderma gangrenosum (PG) is often situated on the lower legs, and the differentiation from conventional leg ulcers (LU) is a challenging task due to the lack of clear clinical diagnostic criteria. Because of the different therapy concepts, misdiagnosis or delayed diagnosis bears a great risk for patients. (2) Objective: to develop a deep convolutional neural network (CNN) capable of analysing wound photographs to facilitate the PG diagnosis for health professionals. (3) Methods: A CNN was trained with 422 expert-selected pictures of PG and LU. In a man vs. machine contest, 33 pictures of PG and 36 pictures of LU were presented for diagnosis to 18 dermatologists at two maximum care hospitals and to the CNN. The results were statistically evaluated in terms of sensitivity, specificity and accuracy for the CNN and for dermatologists with different experience levels. (4) Results: The CNN achieved a sensitivity of 97% (95% confidence interval (CI) 84.2−99.9%) and outperformed dermatologists, with a sensitivity of 72.7% (CI 54.4−86.7%) significantly (p < 0.03). However, dermatologists achieved a slightly higher specificity (88.9% vs. 83.3%). (5) Conclusions: For the first time, a deep neural network was demonstrated to be capable of diagnosing PG, solely on the basis of photographs, and with a greater sensitivity compared to that of dermatologists.
Read moreA Deep Convolution Neural Network Method for Land Cover Mapping: A Case Study of Qinhuangdao, China
Land cover and its dynamic information is the basis for characterizing surface conditions, supporting land resource management and optimization, and assessing the impacts of climate change and human activities. In land cover information extraction, the traditional convolutional neural network (CNN) method has several problems, such as the inability to be applied to multispectral and hyperspectral satellite imagery, the weak generalization ability of the model and the difficulty of automating the construction of a training database. To solve these problems, this study proposes a new type of deep convolutional neural network based on Landsat-8 Operational Land Imager (OLI) imagery. The network integrates cascaded cross-channel parametric pooling and average pooling layer, applies a hierarchical sampling strategy to realize the automatic construction of the training dataset, determines the technical scheme of model-related parameters, and finally performs the automatic classification of remote sensing images. This study used the new type of deep convolutional neural network to extract land cover information from Qinhuangdao City, Hebei Province, and compared the experimental results with those obtained by traditional methods. The results show that: (1) The proposed deep convolutional neural network (DCNN) model can automatically construct the training dataset and classify images. This model performs the classification of multispectral and hyperspectral satellite images using deep neural networks, which improves the generalization ability of the model and simplifies the application of the model. (2) The proposed DCNN model provides the best classification results in the Qinhuangdao area. The overall accuracy of the land cover data obtained is 82.0%, and the kappa coefficient is 0.76. The overall accuracy is improved by 5% and 14% compared to the support vector machine method and the maximum likelihood classification method, respectively.
Read moreDevelopment of a Novel Approach for Classification of MRI Brain Images Using DCNN based onVGG16 Model
Brain tumor implies development of strange cells in brain. In cutting edge stages, brain tumor is most risky infection which can't be relieved. Thus, it ought to be identified in the beginning phases with the assistance of MRI (Magnetic Resonance Image). So, there will be more changes to the patient to endure. Quite possibly, the most functional and significant technique is to utilize Deep Neural Network (DNN). In this paper, a Deep Convolutional Neural Network (DCNN) based on VGG16 model has been developed to identify a tumor through brain Magnetic Resonance Imaging (MRI) dataset. In the clinical field, the strategies of machine learning (ML) and data mining hold a critical stand. It is effectively used to achieve the efficiency and exact location of tumor. The proposed technique involves automatic segmentation method based on Deep Convolution neural network (DCNN). It is layer based segmentation and classification technique. Different levels are engaged with the proposed technique, first step is data collection and then pre-processing & average filtering is done, after that segmentation, feature extraction and classification via DCNN is performed. In this paper, a fresh technique for classification of brain MRI images has been proposed using DCNN based softmax Classifier. The proposed system has been compared with the existing ones. The results are really encouraging as the loss is reduced and accuracy is increased. The proposed system achieved the training accuracy of 98.9% and validation accuracy of 100%. The training loss is reduced up to 0.0230 and the validation loss reduced up to 0.0109.
Read more