• Home
  • Search
  • Integrate template matching and statistical modeling for continuous speech recognition
  • https://doi.org/10.32469/10355/14455Copy DOI Icon

Integrate template matching and statistical modeling for continuous speech recognition

  • Dec 1, 2011
  • Xie Sun
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

In this dissertation, a novel approach of integrating template matching with statistical modeling is proposed to improve continuous speech recognition. Hidden Markov Modeling (HMMs) has been the dominant approach in statistical speech recognition since it provides a principled way of jointly modeling speech spectral variations and time dynamics. However, HMMs have the shortcoming of assuming the observations being independent within each state, which makes it ineffective in modeling the details of speech temporal evolutions that are important for characterizing nonstationary speech sounds. Template-based methods make comparisons between a test pattern and the templates derived from training data, and therefore they are able to capture speech dynamics and time correlation of speech frames better than HMM based methods. However, template matching requires large memory space and computational time since feature vectors of training data need to be stored in computer memory for access at the recognition stage, which is difficult in large vocabulary continuous speech recognition (LVCSR). Our proposed approach takes advantages of both statistical modeling and template matching, which overcomes the weakness of conventional template-based method and is feasible for LVCSR. We use multiple Gaussian Mixture Model (GMM) indices to represent each frame of speech templates, and define the template unit to be context-dependent phone segments (triphone context). We also use phonetic decision trees borrowed from those commonly used in HMMs to tie triphone templates and predict triphones unseen in training data. Two local distances, log likelihood ratio (LLR) and Kullback-Leibler (KL) divergence, are proposed for dynamic time warping (DTW) based template matching. In order to reduce computational complexity and storage space, we propose methods of minimum distance template selection (MDTS) and maximum log-likelihood template selection (MLTS), and investigate a template compression method on top of template selection to further improve recognition performance. The template based methods were used to rescore lattices generated by baseline HMMs on the tasks of TIMIT continuous phone recognition and teleheath LVCSR and experimental results demonstrated that the proposed approach of integrating template matching with statistical modeling significantly improved recognition performances over the HMM baselines. The template selection methods also provided significant recognition accuracy improvements over the HMM baseline while largely reducing the computation and storage complexities. When all templates or MDTS were used, using the LLR local distance obtained better recognition performance than the KL divergence local distance. For MLTS and template compression, KL divergence local distance provided better performance than the LLR local distance, and the template compression method made further improvements over KL based MLTS. Since the templates were constructed based on the GMM indices extracted from HMM baselines, we also validated the effectiveness of the proposed template methods based on enhanced HMM baselines. Experimental results showed that LLR based all template method was able to consistently improve TIMIT phone recognition accuracies based on four enhanced HMM baselines. Prosodic features such as duration, energy, and pitch can reflect longer span information of speech than conventional single frame vectors but they have commonly been ignored by HMMs. Template based methods provide possibilities to conveniently integrate prosodic features into speech recognition, which has not been well studied in the past. In this dissertation, we investigate combining template based methods with the speech prosodic features of duration, energy and pitch to further improve speech recognition accuracy. The scores of prosodic information were computed by a GMM based method and a non-parametric method, and the prosodic scores were combined with the acoustic scores in triphone template matching. Experimental results obtained on the telehealth task showed that prosodic information had positive effects on vowel sound recognition.

Similar Papers
  • Book Chapter
  • Citations25

Deep Neural Network Based Continuous Speech Recognition for Serbian Using the Kaldi Toolkit

  • Jan 01, 2015
  • Branislav Popović +4
  • Research Article
  • Citations12

A HYBRID CONTINUOUS SPEECH RECOGNITION SYSTEM USING SEGMENTAL NEURAL NETS WITH HIDDEN MARKOV MODELS

  • Aug 01, 1993
  • International Journal of Pattern Recognition and Artificial Intelligence
  • G Zavaliagkos +3
  • Conference Article
  • Citations3

Semi-continuous segmental probability modeling for continuous speech recognition

  • Oct 16, 2000
  • Jiyong Zhang +3
  • Conference Article

Large Vocabulary Continuous Audio-Visual Speech Recognition

  • Oct 02, 2018
  • George Sterpu
  • Conference Article
  • Citations241

Large vocabulary continuous speech recognition with context-dependent DBN-HMMS

  • May 01, 2011
  • George E Dahl +3
  • Conference Article
  • Citations2

Upper and lower bounds for approximation of the Kullback-Leibler divergence between Hidden Markov models

  • May 01, 2013
  • Haiyang Li +3
  • Research Article

A 168-mW 2.4×-Real-Time 60-kWord Continuous Speech Recognition Processor VLSI

  • Jan 01, 2013
  • IEICE Transactions on Electronics
  • Guangji He +5
  • Research Article
  • Citations174

Template-Based Continuous Speech Recognition

  • May 01, 2007
  • IEEE Transactions on Audio, Speech and Language Processing
  • Mathias De Wachter +5
  • Research Article
  • Citations21

Japanese large-vocabulary continuous-speech recognition using a newspaper corpus and broadcast news

  • Jun 01, 1999
  • Speech Communication
  • Katsutoshi Ohtsuki +6
  • Research Article
  • Citations58

Using tone information in Cantonese continuous speech recognition

  • Mar 01, 2002
  • ACM Transactions on Asian Language Information Processing
  • Tan Lee +3
  • Conference Article
  • Citations13

Recent improvements of the SpeeD Romanian LVCSR system

  • May 01, 2014
  • Horia Cucu +4
  • Conference Article
  • Citations8

Integrating a non-probabilistic grammar into large vocabulary continuous speech recognition

  • Jan 01, 2005
  • R Beutler +2
  • Conference Article
  • Citations4

Minimum Kullback-Leibler distance based multivariate Gaussian feature adaptation for distant-talking speech recognition

  • May 17, 2004
  • Yue Pan +1
  • Research Article
  • Citations1

A large-vocabulary continuous speech recognition system for Hindi

  • Sep 01, 2004
  • IBM Journal of Research and Development
  • Kumarm +2
  • Book Chapter
  • Citations3

Hybrid Approach for Language Identification Oriented to Multilingual Speech Recognition in the Basque Context

  • Jan 01, 2010
  • N Barroso +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.