- Research Article
534
- 10.1016/j.neunet.2014.08.005
Deep Convolutional Neural Networks for Large-scale Speech Tasks
- Sep 16, 2014
- Neural Networks
- Tara N Sainath + 6 more +6
Deep Convolutional Neural Networks for Large-scale Speech Tasks
Recently, an effective segmentation-free approach via deep neural network based hidden Markov model (DNN-HMM) was proposed and successfully applied to offline handwritten Chinese text recognition. In this study, to further improve the modeling capability, we adopt deep convolutional neural networks (DCNN) to calculate the HMM state posteriors. First, on the frame basis, the DCNN-HMM can automatically learn the features from the raw image of the handwritten text line via the convolutional architecture rather than the handcrafted gradient features using in the DNN-HMM. Second, we examine several important factors of DCNN to the recognition performance, namely the kernel size, the number of blocks and convolutional layers. We also improve the language modeling by using more text data and high-order N-gram. Tested on ICDAR 2013 competition task of CASIA-HWDB database, the proposed DCNN-HMM could achieve a character error rate (CER) of 4.07%, yielding a relative CER reduction of 30.8\% over the DNN-HMM approach. To the best of our knowledge, this is the best published result of the segmentation-free approaches. Furthermore, we explain why DCNN-HMM is more effective than DNN-HMM via the visualization of feature learning and the error pattern analysis.
Deep Convolutional Neural Networks for Large-scale Speech Tasks
Deep Convolutional Neural Networks for Large-scale Speech Tasks
Evolutionary pruning of transfer learned deep convolutional neural network for breast cancer diagnosis in digital breast tomosynthesis
Deep learning models are highly parameterized, resulting in difficulty in inference and transfer learning for image recognition tasks. In this work, we propose a layered pathway evolution method to compress a deep convolutional neural network (DCNN) for classification of masses in digital breast tomosynthesis (DBT). The objective is to prune the number of tunable parameters while preserving the classification accuracy. In the first stage transfer learning, 19 632 augmented regions-of-interest (ROIs) from 2454 mass lesions on mammograms were used to train a pre-trained DCNN on ImageNet. In the second stage transfer learning, the DCNN was used as a feature extractor followed by feature selection and random forest classification. The pathway evolution was performed using genetic algorithm in an iterative approach with tournament selection driven by count-preserving crossover and mutation. The second stage was trained with 9120 DBT ROIs from 228 mass lesions using leave-one-case-out cross-validation. The DCNN was reduced by 87% in the number of neurons, 34% in the number of parameters, and 95% in the number of multiply-and-add operations required in the convolutional layers. The test AUC on 89 mass lesions from 94 independent DBT cases before and after pruning were 0.88 and 0.90, respectively, and the difference was not statistically significant (p > 0.05). The proposed DCNN compression approach can reduce the number of required operations by 95% while maintaining the classification performance. The approach can be extended to other deep neural networks and imaging tasks where transfer learning is appropriate.
Read moreDCNN-based Ship Classification using Enhanced Edge Information and Inception Module
The excellent feature extraction ability of deep convolutional neural networks (DCNNs) has been demonstrated in many image processing tasks, by which image classification can achieve high accuracy with only raw input images. However, the specific image features that influence the classification results are not readily determinable and what lies behind the predictions is unclear. This study proposes a method combining the Sobel and Canny operators and an Inception module for ship classification. The Sobel and Canny operators obtain enhanced edge features from the input images. A convolutional layer is replaced with the Inception module, which can automatically select the proper convolution kernel for ship objects in different image regions. The principle is that the high-level features abstracted by the DCNN, and the features obtained by multi-convolution concatenation of the Inception module must ultimately derive from the edge information of the preprocessing input images. This indicates that the classification results are based on the input edge features, which indirectly interpret the classification results to some extent. Experimental results show that the combination of the edge features and the Inception module improves DCNN ship classification performance. The original model with the raw dataset has an average accuracy of 88.72%, while when using enhanced edge features as input, it achieves the best performance of 90.54% among all models. The model that replaces the fifth convolutional layer with the Inception module has the best performance of 89.50%. It performs close to VGG-16 on the raw dataset and is significantly better than other deep neural networks. The results validate the functionality and feasibility of the idea posited.
Read moreAbsolute distance measurement based on laser self-mixing interferometry and deep neural network
Self-mixing interferometry (SMI) is superior to other laser interferometry methods due to its simplicity and compactness. However, SMI signals are often complex and difficult to process due to interference in the form of variations in the effective reflectivity of the target, noisy signals, complex signal shapes, and other dependencies. Deep neural networks have been a very popular area of research in computer artificial intelligence in recent years, allowing more implicit features to be uncovered than traditional shallow machine learning. It has been shown that convolutional and back-propagation neural networks can be used for SMI signal processing. There are also studies that have used machine learning genetic algorithms for absolute distance measurement. Based on the above, this study used convolutional neural networks to form a deep neural network for absolute distance measurement based on SMI technology. We first trained the deep convolutional neural network at different feedback strengths. The results of the Convolutional Neural Network (CNN) model showed a coefficient of determination of 0.9987. which is consistent with the required model. The trained network was then used to estimate absolute distances with and without the addition of noise. The comparison proves that the proposed method is noise-proof and has high adaptability for measurements under different conditions.
Read moreHandwritten text recognition of historical documents using deep neural network technologies
The application of deep neural network technologies to the problem of handwriting recognition in pre-reform Russian is considered. The initial data used are scanned JPG images of historical documents from the 19 th century, in particular containing various noises and interference, which complicates the work of the recognition algorithm. Text recognition is performed in three stages: noise removal, segmentation (highlighting) of text lines in the image, since the input data for the deep neural network are precisely the lines, and then recognition of the text of the highlighted lines using the pre-trained Tesseract OCR model, which performs electronic translation of images of handwritten or printed text into text data. The model used is a convolutional recurrent neural network; the model is a combination of a convolutional neural network for extracting local features from an image and a recurrent neural network represented by two layers of bidirectional LSTM networks for processing the sequence. Using this model allows for reliable recognition of handwritten text.
Read moreClassifying multi-category images using deep learning : A convolutional neural network model
This paper presents an image classification model using a convolutional neural network with Tensor Flow. Tensor Flow is a popular open source library for machine learning and deep neural networks. A multi-category image dataset has been considered for the classification. Conventional back propagation neural network has an input layer, hidden layer, and an output layer but convolutional neural network, has a convolutional layer, and a max pooling layer. We train this proposed classifier to calculate the decision boundary of the image dataset. The data in the real world is mostly in the form of unlabeled and unstructured format. These unstructured data may be image, sound and text data. Useful information cannot be easily derived from neural networks which are shallow i.e. the ones which have less number of hidden layers. We propose deep neural network based CNN classifier which has a large number of hidden layers and can derive meaningful information from images.
Read moreHandwritten Digit Recognition Using Very Deep Convolutional Neural Network
Automated image classification is an essential task of the computer vision field. The tagging of images into a set of predefined groups is referred to as image classification. The implementation of computer vision to automate image classification would be beneficial because manual image evaluation and identification can be time-consuming, particularly when there are many images of different classes. Deep learning approaches are proven to overperform existing machine learning techniques in many fields in recent years, and computer vision is one of the most notable examples. The very deep neural network (VDCNN) is a powerful deep learning model for image classification, and this paper examines it briefly using MNIST handwritten digit dataset. This dataset is used to prove the efficacy of very deep neural networks over other deep learning models. The proposed study aims to comprehend the very deep neural network architecture used to accomplish a handwritten digit recognition task. The feasibility of the proposed model is evaluated using mean accuracy, validation accuracy, and standard deviation. The study results of the very deep neural network model are compared to a convolutional neural network and convolutional neural network with batch normalization. According to the results of the comparison study, very deep neural networks achieve high accuracy of 99.1% for a handwritten dataset. The outcome of the proposed work is used to interpret how well a very deep neural network performs when compared to the other two models of deep neural networks. This proposed architecture may be used to automate the classification of handwritten digits dataset.
Read moreMultimodal and Crossmodal Representation Learning from Textual and Visual Features with Bidirectional Deep Neural Networks for Video Hyperlinking
Video hyperlinking represents a classical example of multimodal problems. Common approaches to such problems are early fusion of the initial modalities and crossmodal translation from one modality to the other. Recently, deep neural networks, especially deep autoencoders, have proven promising both for crossmodal translation and for early fusion via multimodal embedding. A particular architecture, bidirectional symmetrical deep neural networks, have been proven to yield improved multimodal embeddings over classical autoencoders, while also being able to perform crossmodal translation. In this work, we focus firstly at evaluating good single-modal continuous representations both for textual and for visual information. Word2Vec and paragraph vectors are evaluated for representing collections of words, such as parts of automatic transcripts and multiple visual concepts, while different deep convolutional neural networks are evaluated for directly embedding visual information, avoiding the creation of visual concepts. Secondly, we evaluate methods for multimodal fusion and crossmodal translation, with different single-modal pairs, in the task of video hyperlinking. Bidirectional (symmetrical) deep neural networks were shown to successfully tackle downsides of multimodal autoencoders and yield a superior multimodal representation. In this work, we extensively tests them in different settings, with different single-modal representations, within the context of video-hyperlinking. Our novel bidirectional symmetrical deep neural networks are compared to classical autoencoders and are shown to yield significantly improved multimodal embeddings that significantly (alpha=0.0001) outperform multimodal embeddings obtained by deep autoencoders with an absolute improvement in precision at 10 of 14.1% when embedding visual concepts and automatic transcripts and an absolute improvement of 4.3% when embedding automatic transcripts with features obtained with very deep convolutional neural networks, yielding 80% of precision at 10.
Read moreConvolutional neural network-based wind pressure prediction on low-rise buildings
Convolutional neural network-based wind pressure prediction on low-rise buildings
Maxout neurons for deep convolutional and LSTM neural networks in speech recognition
Maxout neurons for deep convolutional and LSTM neural networks in speech recognition
DEEP LEARNING FRAMEWORK FOR WOVEN COMPOSITE ANALYSIS
paper, we focus on exploring the relationship between weave patterns and their mechanical properties in woven fiber composites through Machine Learning. Specifically, we explore the interactions between woven architectures and in-plane stiffness properties through Deep Convolutional Neural Network (DCNN) and Generative Adversarial Network (GAN). Our research is important for exploring how woven composite’s pattern is related to its mechanical properties and accelerating woven composite design as well as optimization. We focus on two tasks: (1) Stiffness prediction: Predicting in-plane stiffness properties for given weave patterns. Our DCNN extracts high-level features through several convolutional and fully connected layers to determine the final predictions. (2) Weave pattern prediction: Predicting weave patterns for target stiffness properties, which can be treated as the reverse task of the first one. Due to many-to-one mapping between weave patterns and the composite properties, we utilize a Decoder Neural Network as our baseline model and compare its performance with GAN and Genetic Algorithm. We represent the weave patterns as 2D checkerboard models and use finite element analysis (FEA) to determine in-plane stiffness properties, which serve as input data for our ML framework. We show that: (1) for stiffness prediction, DCNN can predict stiffness values for a given weave pattern with relatively high accuracy (above 93%); (2) for weave pattern prediction, the GAN model gives the best prediction accuracy (above 92%) while Decoder Neural Network has the best time efficiency. HAOTIAN FENG
Read moreDeepArrNet: An Efficient Deep CNN Architecture for Automatic Arrhythmia Detection and Classification From Denoised ECG Beats
In this paper, an efficient deep convolutional neural network (CNN) architecture is proposed based on depthwise temporal convolution along with a robust end-to-end scheme to automatically detect and classify arrhythmia from denoised electrocardiogram (ECG) signal, which is termed as `DeepArrNet'. Firstly, considering the variational pattern of wavelet denoised ECG data, a realistic augmentation scheme is designed that offers a reduction in class imbalance as well as increased data variations. A structural unit, namely PTP (Pontwise-Temporal-Pointwise Convolution) unit, is designed with its variants where depthwise temporal convolutions with varying kernel sizes are incorporated along with prior and post pointwise convolution. Afterward, a deep neural network architecture is constructed based on the proposed structural unit where series of such structural units are stacked together while increasing the kernel sizes for depthwise temporal convolutions in successive units along with the residual linkage between units through feature addition. Moreover, multiple depthwise temporal convolutions are introduced with varying kernel sizes in each structural unit to make the process more efficient while strided convolutions are utilized in the residual linkage between subsequent units to compensate the increased computational complexity. This architecture provides the opportunity to explore the temporal features in between convolutional layers more optimally from different perspectives utilizing diversified temporal kernels. Extensive experimentations are carried out on two publicly available datasets to validate the proposed scheme that results in outstanding performances in all traditional evaluation metrics outperforming other state-of-the-art approaches.
Read moreDeep Learning Analysis in Development of Handwritten and Plain Text Classification API
Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) are technologies that enable text recognition. The difference between OCR and HTR is one designed specifically for digital text and one designed for handwritten text. There are already various implementations of OCR and HTR online. However, such systems do not guarantee the systems are in premises. To solve this problem, the OCR and HTR system must be built from the scratch. The purpose of this research is to improve the recognition by separating the text whether it is a handwritten or a printed text, which will later be forwarded into the appropriate recognition system. An application program interface (API) was also created in order to finalize the classification system into real world usage. In this research, the classification system being developed using convolutional neural network (CNN) method. To be able to reach the highest accuracy of the classification system, the experimentation and improvement about hyperparameters, dataset format, data augmentation and analysis on 3 CNN architectures were conducted. In the end of this research, there are 2 architectures in a tight competition, one is VGG-16 with 90.63% accuracy and one is AlexNet with 90.17% accuracy on ideal data testing. However, AlexNet is chosen as the winner after the testing with real data.
Read moreEvaluation of the benchmark datasets for testing the efficacy of deep convolutional neural networks
Evaluation of the benchmark datasets for testing the efficacy of deep convolutional neural networks
Adversarial Robustness of Deep Convolutional Neural Network-based Image Recognition Models: A Review
Deep convolutional neural networks have achieved great success in recent years. They have been widely used in various applications such as optical and SAR image scene classification, object detection and recognition, semantic segmentation, and change detection. However, deep neural networks rely on large-scale high-quality training data, and can only guarantee good performance when the training and test data are independently sampled from the same distribution. Deep convolutional neural networks are found to be vulnerable to subtle adversarial perturbations. This adversarial vulnerability prevents the deployment of deep neural networks in security-sensitive applications such as medical, surveillance, autonomous driving and military scenarios. This paper first presents a holistic view of security issues for deep convolutional neural network-based image recognition systems. The entire information processing chain is analyzed regarding safety and security risks. In particular, poisoning attacks and evasion attacks on deep convolutional neural networks are analyzed in detail. The root causes of adversarial vulnerabilities of deep recognition models are also discussed. Then, we give a formal definition of adversarial robustness and present a comprehensive review of adversarial attacks, adversarial defense, and adversarial robustness evaluation. Rather than listing existing research, we focus on the threat models for the adversarial attack and defense arms race. We perform a detailed analysis of several representative adversarial attacks on SAR image recognition models and provide an example of adversarial robustness evaluation. Finally, several open questions are discussed regarding recent research progress from our workgroup. This paper can be further used as a reference to develop more robust deep neural network-based image recognition models in dynamic adversarial scenarios.
Read more