- Research Article
58
- 10.1016/j.csl.2004.12.002
Product of Gaussians for speech recognition
- Jan 26, 2005
- Computer Speech & Language
- M.J.F Gales + 1 more +1
Product of Gaussians for speech recognition
Recently, deep neural network (DNN) with hidden Markov model (HMM) has turned out to be a superior sequence learning framework, based on which significant improvements were achieved in many application tasks, such as automatic speech recognition (ASR). However, the training of DNN-HMM requires the pre-segmented training data, which can be generated using Gaussian Mixture Model (GMM) in ASR tasks. Thus, questions are raised by many researchers: can we train the DNN-HMM without GMM seeding, and what does it suggest if the answer is yes? In this research, we come up with the ‘yes’ answer by presenting forward-backward learning algorithm for DNN-HMM framework. Besides, a training procedure is proposed, in which, the training for context independent (CI) DNN-HMM is treated as the pre-training for context dependent (CD) DNN-HMM. To evaluate the contribution of this work, experiments on ASR task with the benchmark corpus TIMIT are performed, and the results demonstrate the effectiveness of this research.
Product of Gaussians for speech recognition
Product of Gaussians for speech recognition
A case study of speech recognition in Spanish: From conventional to deep approach
The aim of this paper is to exhibit a comparative case study of the conventional speech recognition GMM-HMM (Gaussian mixture model - hidden Markov model) architecture and the recent model based on deep neural networks. During years the GMM approach has controlled the speech recognition tasks, however it has been surpassed with the resurgence of artificial neural networks. To exemplify these acoustic modeling frameworks, a case study has been conducted by using the Kaldi toolkit, employing a personalized speaker-independent mid-vocabulary voice corpus for recognition of digit strings and personal name lists in latin spanish on a connected-words phone dialing task. The speech recognition accuracy obtained in the results shows a better word error rate by using the DNN acoustic modeling. A 20.71% relative improvement is obtained with DNN-HMM models (3.33% WER) in respect to the lowest GMM-HMM rate (4.20% WER).
Read moreFMLLR Speaker Normalization With i-Vector: In Pseudo-FMLLR and Distillation Framework
When an automatic speech recognition (ASR) system is deployed for real-world applications, it often receives only one utterance at a time for decoding. This single utterance could be of short duration depending on the ASR task. In these cases, robust estimation of speaker normalizing methods like feature-space maximum likelihood linear regression (FMLLR) and i -vectors may not be feasible. In this paper, we propose two unsupervised speaker normalization techniques—one at feature level and other at model level of acoustic modeling—to overcome the drawbacks of FMLLR and i -vectors in real-time scenarios. At feature level, we propose the use of deep neural networks (DNN) to generate pseudo-FMLLR features from time-synchronous pair of filterbank and FMLLR features. These pseudo-FMLLR features can then be used for DNN acoustic model training and decoding. At model level, we propose a generalized distillation framework, where a teacher DNN trained on FMLLR features guides the training and optimization of a student DNN trained on filterbank features. In both the proposed methods, the ambiguity in choosing the speaker-specific FMLLR transform can be reduced by augmenting i -vectors to the input filterbank features. Experiments conducted on 33-h and 110-h subsets of Switchboard corpus show that the proposed methods provide significant gains over DNNs trained on FMLLR, i -vector appended FMLLR, filterbank and i -vector appended filterbank features, in real-time scenario.
Read moreSpeech Recognition in Indian Languages—A Survey
In this paper, a brief overview derived out of detailed survey of speech recognition works reported in Indian languages is described. Robustness of speech recognition systems toward language variation is the recent trend of research in speech recognition technology. To develop a system which can communicate with human in any language like any other human is the foremost requirement in order to design appropriate speech recognition technology for one to all. India is a country which has vast linguistic variations among its billion plus population. Therefore, it provides a sound area of research toward language-specific speech recognition technology. From the beginning of the commercial availability of the speech recognition system, the technology has been dominated by the hidden Markov model (HMM) methodology due to its capability of modeling temporal structures of speech and encoding them as a sequence of spectral vectors. Most of the work done in Indian languages also uses HMM technology. However, from the last 10–15 years after the acceptance of neurocomputing as an alternative to HMM, artificial neural network (ANN)-based methodologies have started to receive attention for application in speech recognition. This is a trend worldwide as part of which few works have also been reported by a few researchers.
Read moreAn Information-Extraction Approach to Speech Processing: Analysis, Detection, Verification, and Recognition
The field of automatic speech recognition (ASR) has enjoyed more than 30 years of technology advances due to the extensive utilization of the hidden Markov model (HMM) framework and a concentrated effort by the speech community to make available a vast amount of speech and language resources, known today as the Big Data Paradigm. State-of-the-art ASR systems achieve a high recognition accuracy for well-formed utterances of a variety of languages by decoding speech into the most likely sequence of words among all possible sentences represented by a finite-state network (FSN) approximation of all the knowledge sources required by the ASR task. However, the ASR problem is still far from being solved because not all information available in the speech knowledge hierarchy can be directly integrated into the FSN to improve the ASR performance and enhance system robustness. It is believed that some of the current issues of integrating various knowledge sources in top-down integrated search can be partially addressed by processing techniques that take advantage of the full set of acoustic and language information in speech. It has long been postulated that human speech recognition (HSR) determines the linguistic identity of a sound based on detected evidence that exists at various levels of the speech knowledge hierarchy, ranging from acoustic phonetics to syntax and semantics. This calls for a bottom-up attribute detection and knowledge integration framework that links speech processing with information extraction, by spotting speech cues with a bank of attribute detectors, weighting and combining acoustic evidence to form cognitive hypotheses, and verifying these theories until a consistent recognition decision can be reached. The recently proposed automatic speech attribute transcription (ASAT) framework is an attempt to mimic some HSR capabilities with asynchronous speech event detection followed by bottom-up knowledge integration and verification. In the last few years, ASAT has demonstrated good potential and has been applied to a variety of existing applications in speech processing and information extraction.
Read moreP‐3.3: Deep Learning in Speech Recognition: A Review of Recent Methods, Applications and Advantages
In recent years, with the rapid development of Deep Learning (DL) and the widespread uses of Deep Neural Networks (DNN), speech recognition technology has attracted great attention. This technology enables machines to convert human speech into text and plays an important role in some areas such as virtual assistants, conference transcription and smart home devices. Traditional models such as Hidden Markov Models (HMM) and Gaussian Mixture Models (GMM) used to play a key role in modeling and recognizing speech signals. Even though these methods still have challenges in dealing with high‐dimensional feature vectors as they struggle to efficiently learn large amounts of model parameters while maintaining good generalization performance. The emergence of DNNs has revolutionized the field of speech recognition, making it possible to automatically extract features from raw speech signals. Thereby, modern speech recognition systems can exhibit higher adaptability and generalization capabilities, achieving higher accuracy and stability under various conditions. This review explores the application of DL techniques to speech recognition, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformer‐based architectures, etc. Advances in these techniques have dramatically improved recognition accuracy, noise immunity, and adaptability to multiple accents, thereby improving the total efficiency of automatic speech systems. By leveraging advanced feature extraction and modeling capabilities, deep learning‐based systems support real‐ time processing and minimize the dependency on traditional manually designed features. However, there are still limitations in terms of interpretability, computational efficiency, and fairness to different accents and languages, which impede the widespread use of these techniques. Future research should focus on solving these issues in order to achieve fairer, efficient, and scalable deep learning‐driven speech recognition schemes.
Read moreComplex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition
The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with different scales of neuronal dynamics is necessary. Here we introduce four types of neuronal dynamics to post-process the sequential patterns generated from the spiking transformer to get the complex dynamic neuron improved spiking transformer neural network (DyTr-SNN). We found that the DyTr-SNN could handle the non-toy automatic speech recognition task well, representing a lower phoneme error rate, lower computational cost, and higher robustness. These results indicate that the further cooperation of SNNs and neural dynamics at the neuron and network scales might have much in store for the future, especially on the ASR tasks.
Read moreIntegrate template matching and statistical modeling for continuous speech recognition
In this dissertation, a novel approach of integrating template matching with statistical modeling is proposed to improve continuous speech recognition. Hidden Markov Modeling (HMMs) has been the dominant approach in statistical speech recognition since it provides a principled way of jointly modeling speech spectral variations and time dynamics. However, HMMs have the shortcoming of assuming the observations being independent within each state, which makes it ineffective in modeling the details of speech temporal evolutions that are important for characterizing nonstationary speech sounds. Template-based methods make comparisons between a test pattern and the templates derived from training data, and therefore they are able to capture speech dynamics and time correlation of speech frames better than HMM based methods. However, template matching requires large memory space and computational time since feature vectors of training data need to be stored in computer memory for access at the recognition stage, which is difficult in large vocabulary continuous speech recognition (LVCSR). Our proposed approach takes advantages of both statistical modeling and template matching, which overcomes the weakness of conventional template-based method and is feasible for LVCSR. We use multiple Gaussian Mixture Model (GMM) indices to represent each frame of speech templates, and define the template unit to be context-dependent phone segments (triphone context). We also use phonetic decision trees borrowed from those commonly used in HMMs to tie triphone templates and predict triphones unseen in training data. Two local distances, log likelihood ratio (LLR) and Kullback-Leibler (KL) divergence, are proposed for dynamic time warping (DTW) based template matching. In order to reduce computational complexity and storage space, we propose methods of minimum distance template selection (MDTS) and maximum log-likelihood template selection (MLTS), and investigate a template compression method on top of template selection to further improve recognition performance. The template based methods were used to rescore lattices generated by baseline HMMs on the tasks of TIMIT continuous phone recognition and teleheath LVCSR and experimental results demonstrated that the proposed approach of integrating template matching with statistical modeling significantly improved recognition performances over the HMM baselines. The template selection methods also provided significant recognition accuracy improvements over the HMM baseline while largely reducing the computation and storage complexities. When all templates or MDTS were used, using the LLR local distance obtained better recognition performance than the KL divergence local distance. For MLTS and template compression, KL divergence local distance provided better performance than the LLR local distance, and the template compression method made further improvements over KL based MLTS. Since the templates were constructed based on the GMM indices extracted from HMM baselines, we also validated the effectiveness of the proposed template methods based on enhanced HMM baselines. Experimental results showed that LLR based all template method was able to consistently improve TIMIT phone recognition accuracies based on four enhanced HMM baselines. Prosodic features such as duration, energy, and pitch can reflect longer span information of speech than conventional single frame vectors but they have commonly been ignored by HMMs. Template based methods provide possibilities to conveniently integrate prosodic features into speech recognition, which has not been well studied in the past. In this dissertation, we investigate combining template based methods with the speech prosodic features of duration, energy and pitch to further improve speech recognition accuracy. The scores of prosodic information were computed by a GMM based method and a non-parametric method, and the prosodic scores were combined with the acoustic scores in triphone template matching. Experimental results obtained on the telehealth task showed that prosodic information had positive effects on vowel sound recognition.
Read moreDeep Neural Network Based Continuous Speech Recognition for Serbian Using the Kaldi Toolkit
This paper presents a deep neural network (DNN) based large vocabulary continuous speech recognition (LVCSR) system for Serbian, developed using the open-source Kaldi speech recognition toolkit. The DNNs are initialized using stacked restricted Boltzmann machines (RBMs) and trained using cross-entropy as the objective function and the standard error backpropagation procedure in order to provide posterior probability estimates for the hidden Markov model (HMM) states. Emission densities of HMM states are represented as Gaussian mixture models (GMMs). The recipes were modified based on the particularities of the Serbian language in order to achieve the optimal results. A corpus of approximately 90 hours of speech (21000 utterances) is used for the training. The performances are compared for two different sets of utterances between the baseline GMM-HMM algorithm and various DNN settings.KeywordsKaldi speech recognition toolkitContinuous speech recognitionDeep neural networksSerbian
Read moreDialect Identification in Telugu Language Speech Utterance Using Modified Features with Deep Neural Network
Dialect Identification is the process of identifies the dialects of particular standard language. The Telugu Language is one of the historical and important languages. Like any other language Telugu also contains mainly three dialects Telangana, Costa Andhra and Rayalaseema. The research work in dialect identification is very less compare to Language identification because of dearth of database. In any dialects identification system, the database and feature engineering play vital roles because of most the words are similar in pronunciation and also most of the researchers apply statistical approaches like Hidden Markov Model (HMM), Gaussian Mixture Model (GMM), etc. to work on speech processing applications. But in today's world, neural networks play a vital role in all application domains and produce good results. One of the types of the neural networks is Deep Neural Networks (DNN) and it is used to achieve the state of the art performance in several fields such as speech recognition, speaker identification. In this, the Deep Neural Network (DNN) based model Multilayer Perceptron is used to identify the regional dialects of the Telugu Language using enhanced Mel Frequency Cepstral Coefficients (MFCC) features. To do this, created a database of the Telugu dialects with the duration of 5h and 45m collected from different speakers in different environments. The results produced by DNN model compared with HMM and GMM model and it is observed that the DNN model provides good performance.
Read moreNon-Uniform Boosted MCE Training of Deep Neural Networks for Keyword Spotting
Keyword spotting can be formulated as a non-uniform error automatic speech recognition (ASR) problem. It has been demonstrated [1] that this new formulation with the nonuniform MCE training technique can lead to improved system performance in keyword spotting applications. In this paper, we demonstrate that deep neural networks (DNNs) can be successfully trained on the non-uniform minimum classification error (MCE) criterion which weighs the errors on keywords much more significantly than those on non-keywords in an ASR task. The integration with a DNN-HMM system enables modeling of multi-frame distributions, which conventional systems find difficult to accomplish. To further improve the performance, more confusable data is generated by boosting the likelihood of the sentences that have more errors. The keyword spotting system is implemented within a weighted finite state transducer (WFST) framework and the DNN is optimized using standard backpropagation and stochastic gradient decent. We evaluate the performance of the proposed framework on a large vocabulary spontaneous conversational telephone speech dataset (Switchboard-1 Release 2). The proposed approach achieves an absolute figure of merit improvement of 3.65% over the baseline system.
Read moreThe Effect of Different Optimization Techniques on End-to-End Turkish Speech Recognition Systems that use Connectionist Temporal Classification
In the production of acoustic models for speech recognition applications, the use of Long Short Term Memory(LSTM) based Recurrent Neural Network(RNN) has begun to get better results than the use of Gaussian Mixture Model(GMM). The creation of GMM-based acoustic models is prolonging the deep learning process due to the need for aligned Hidden Markov Model(HMM). As a solution to this problem, another method to generate acoustic models is proposed that is based on Connectionist Temporal Classification(CTC). In this study, a CTC based model is created and the effect of different optimization techniques on the classification performance is compared. These tests were applied on Turkish speech datasets to determine the best optimization techniques to be used in speech recognition applications. Our evaluation results showed that GradientDescent, ProximalGradientDescent and RMSPROP produce better results than other algorithms.
Read moreFeature mapping using far-field microphones for distant speech recognition
Feature mapping using far-field microphones for distant speech recognition
Speaker Identity Recognition by Acoustic and Visual Data Fusion through Personal Privacy for Smart Care and Service Applications
With rapid developments in techniques related to the internet of things, smart service applications such as voice-command-based speech recognition and smart care applications such as context-aware-based emotion recognition will gain much attention and potentially be a requirement in smart home or office environments. In such intelligence applications, identity recognition of the specific member in indoor spaces will be a crucial issue. In this study, a combined audio-visual identity recognition approach was developed. In this approach, visual information obtained from face detection was incorporated into acoustic Gaussian likelihood calculations for constructing speaker classification trees to significantly enhance the Gaussian mixture model (GMM)-based speaker recognition method. This study considered the privacy of the monitored person and reduced the degree of surveillance. Moreover, the popular Kinect sensor device containing a microphone array was adopted to obtain acoustic voice data from the person. The proposed audio-visual identity recognition approach deploys only two cameras in a specific indoor space for conveniently performing face detection and quickly determining the total number of people in the specific space. Such information pertaining to the number of people in the indoor space obtained using face detection was utilized to effectively regulate the accurate GMM speaker classification tree design. Two face-detection-regulated speaker classification tree schemes are presented for the GMM speaker recognition method in this study—the binary speaker classification tree (GMM-BT) and the non-binary speaker classification tree (GMM-NBT). The proposed GMM-BT and GMM-NBT methods achieve excellent identity recognition rates of 84.28% and 83%, respectively; both values are higher than the rate of the conventional GMM approach (80.5%). Moreover, as the extremely complex calculations of face recognition in general audio-visual speaker recognition tasks are not required, the proposed approach is rapid and efficient with only a slight increment of 0.051 s in the average recognition time.
Read moreArabic phonemes recognition system based on malay speakers using neural network
Arabic language can be used by native and non- native speakers; due to Arabic is the language of the holy book of Muslims. In this paper, Arabic phoneme recognition system is proposed based on Malay speakers. This system consists of three main stages. The first stage is noise reduction and it aims to enhance the phoneme signals by excluding the unvoiced signals and keep only the voiced signal. Wiener filter is adapted to accomplish this task. The second stage is based on Mel-Frequency Cepstral Coefficients method to extract a vector of features to represent each phoneme signal. Eventually, pattern recognition neural network is designed as recognizer. The proposed system produces sufficient outcomes with 20 hidden neurons. Keywords—Arabic; Malay; pattern recognition; Wiener; Mel- Frequency Cepstral Coefficients I. INTRODUCTION It is well known the wide range of Artificial Neural Networks (ANN) application in speech recognition, financial, telecommunications, electronics and medical. As for speech recognition system for Arabic language, (1) & (2) have addressed the implementation of recognition system based on Arabic native speakers. In general, neural networks (NNs) for phonemes recognition are divided into several categories for instance Very Large Scale Integration (VLSI) NN that can be divided into digital and analogue with digital NNs being more compatible with feed-forward neural networks while analogue NNs are found to be more successful with recurrent NNs (3). From Automatic Speech System (ASR) point of view, NNs for recognition purpose can be divided into two main category namely conventional neural networks such as Multilayer Perceptron (MLP) and Radial Basis Function (RBF) and secondly Recurrent Neural Networks (RNN). The first category of NNs was implemented as pattern classifiers and proven capable for recognition purpose but not at par as compared to Hidden Markov Model (HMM) performance (2). At present, HMM has been proven to be the most successful approach in ASR research area until recently the possibility to hybrid both HMM and ANN (4) & (5). In general, speech recognition systems consist of three stages specifically noise reduction, feature extraction and classification. The aim of noise reduction is to reduce the level of noise in the speech signal and make it more compatible for the ASR. As for feature extraction, this process attempted to extract the most relevant features related to the phoneme signal. The extracted features acted as input features of the NN during training and testing. Finally, the purpose of classification process is to arrange the most similar and related features in categories based on the training phase. Basically, networks training database are being observed and updated during training session of the networks, involving the weights and biases arguments. Networks with lower errors can be considered as better networks (6). Hence, this paper deems to explore further the ability of NN for Arabic phonemes recognition spoken by Malay individuals. This paper is arranged as following; related work is discussed in section II followed by section III that discussed method to be implemented based on Wiener filter, Mel-Frequency Cepstral Coefficients and pattern recognition neural network. Section IV elaborated experimental analysis and results and followed by Conclusion section.
Read more