• Home
  • Search
  • Incorporating language constraints in sub-word based speech recognition
  • Cite Icon42
  • https://doi.org/10.1109/asru.2005.1566516Copy DOI Icon

Incorporating language constraints in sub-word based speech recognition

  • Jan 1, 2005
  • H Erdogan +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In large vocabulary continuous speech recognition (LVCSR) for agglutinative and inflectional languages, we encounter problems due to theoretically infinite full-word lexicon size. Sub-word lexicon units may be utilized to dramatically reduce the out-of-vocabulary rate in test data. One can develop language models based on sub-word units to perform LVCSR. However, it has not always been beneficial to use sub-word lexicon units, since shorter units have higher acoustic confusability among them and language model history is effectively shorter as compared to the history in full-word language models. To reduce the aforementioned problems, we propose using the longest possible sub-word units in our lexicon, namely half-words and full-words only. We also incorporate linguistic rules of half-word combination into our statistical language model. The language constraints are represented with a rule-based WFSM which can be combined with an N-gram language model to yield a better and smaller language model. We study the performance of the proposed system for Turkish LVCSR, when the language constraint takes the form of enforcing vowel harmony between stems and endings. We also introduce novel error-rate metrics that are more appropriate than word-error-rate for agglutinative languages. Using half-words with a bi-gram model yields a significant reduction in word-error-rate as compared to a bi-gram full-word model. In addition, combining a tri-gram half-word language model with the vowel-harmony WFSM improves the accuracy further when rescoring the bi-gram lattices.

Similar Papers
  • Research Article
  • Citations18

A fast and memory-efficient N-gram language model lookup method for large vocabulary continuous speech recognition

  • Dec 19, 2005
  • Computer Speech & Language
  • Xiaolong Li +1
  • Research Article
  • Citations4

Construction and evaluation of language models based on stochastic context‐free grammar for speech recognition

  • Oct 23, 2002
  • Systems and Computers in Japan
  • Chiori Hori +3
  • Research Article
  • Citations6

USING DATA-DRIVEN SUBWORD UNITS IN LANGUAGE MODEL OF HIGHLY INFLECTIVE SLOVENIAN LANGUAGE

  • Mar 01, 2009
  • International Journal of Pattern Recognition and Artificial Intelligence
  • Mirjam Sepesy Maučec +3
  • Research Article
  • Citations21

Japanese large-vocabulary continuous-speech recognition using a newspaper corpus and broadcast news

  • Jun 01, 1999
  • Speech Communication
  • Katsutoshi Ohtsuki +6
  • Book Chapter
  • Citations4

Comparison of Grapheme and Phoneme Based Acoustic Modeling in LVCSR Task in Slovak

  • Jan 01, 2009
  • Michal Mirilovič +2
  • Conference Article
  • Citations14

A phone-viseme dynamic Bayesian network for audio-visual automatic speech recognition

  • Dec 01, 2008
  • Louis Terry +1
  • Research Article
  • Citations12

Subspace Gaussian mixture based language modeling for large vocabulary continuous speech recognition

  • Jan 23, 2020
  • Speech Communication
  • Ri Hyon Sun +1
  • Conference Article
  • Citations4

An LVCSR Based Automatic Scoring Method in English Reading Tests

  • Aug 01, 2012
  • Junbo Zhang +2
  • Research Article
  • Citations10

An improved two-stage mixed language model approach for handling out-of-vocabulary words in large vocabulary continuous speech recognition

  • Apr 17, 2013
  • Computer Speech & Language
  • Bert Réveil +2
  • Conference Article
  • Citations20

Developing STT and KWS systems using limited language resources

  • Sep 14, 2014
  • Viet-Bac Le +7
  • Research Article
  • Citations22

Modelling Semantic Context of OOV Words in Large Vocabulary Continuous Speech Recognition

  • Feb 08, 2017
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Imran Sheikh +3
  • Conference Article
  • Citations372

Improved Training of End-to-end Attention Models for Speech Recognition

  • Sep 02, 2018
  • Albert Zeyer +3
  • Conference Article
  • Citations13

Recent improvements of the SpeeD Romanian LVCSR system

  • May 01, 2014
  • Horia Cucu +4
  • Conference Article

Large Vocabulary Continuous Audio-Visual Speech Recognition

  • Oct 02, 2018
  • George Sterpu
  • Conference Article
  • Citations5

A Multi-Genre Urdu Broadcast Speech Recognition System

  • Nov 18, 2021
  • Erbaz Khan +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.