• Home
  • Search
  • Transliteration Based Approaches to Improve Code-Switched Speech Recognition Performance
  • Cite Icon41
  • https://doi.org/10.1109/slt.2018.8639699Copy DOI Icon

Transliteration Based Approaches to Improve Code-Switched Speech Recognition Performance

  • Dec 1, 2018
  • Jesse Emond +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Code-switching is a commonly occurring phenomenon in many multilingual communities, wherein a speaker switches between languages within a single utterance. Conventional Word Error Rate (WER) is not sufficient for measuring the performance of code-mixed languages due to ambiguities in transcription, misspellings and borrowing of words from two different writing systems. These rendering errors artificially inflate the WER of an Automated Speech Recognition (ASR) system and complicate its evaluation. Furthermore, these errors make it harder to accurately evaluate modeling errors originating from code-switched language and acoustic models. In this work, we propose the use of a new metric, transliteration-optimized Word Error Rate (toWER) that smoothes out many of these irregularities by mapping all text to one writing system and demonstrate a correlation with the amount of code-switching present in a language. We also present a novel approach to acoustic and language modeling for bilingual code-switched Indic languages using the same transliteration approach to normalize the data for three types of language models, namely, a conventional n-gram language model, a maximum entropy based language model and a Long Short Term Memory (LSTM) language model, and a state-of-the-art Connectionist Temporal Classification (CTC) acoustic model. We demonstrate the robustness of the proposed approach on several Indic languages from Google Voice Search traffic with significant gains in ASR performance up to 10% relative over the state-of-the-art baseline.

Similar Papers
  • Research Article
  • Citations6

Croatian Large Vocabulary Automatic Speech Recognition

  • Jan 18, 2017
  • Automatika
  • Sanda Martinčić-Ipšić +2
  • Research Article
  • Citations3

Armenian Speech Recognition System: Acoustic and Language Models

  • Jan 01, 2022
  • International Journal Of Scientific Advances
  • Varuzhan H Baghdasaryan
  • Research Article
  • Citations41

Multi-microphone speech recognition integrating beamforming, robust feature extraction, and advanced DNN/RNN backend

  • Feb 27, 2017
  • Computer Speech & Language
  • Takaaki Hori +6
  • Dissertation
  • Citations1

An acoustic study of syllable rhymes : a basis for Thai continuous speech recognition system

  • Jan 01, 2003
  • Ekkarit Maneenoi
  • Conference Article
  • Citations54

Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition

  • May 01, 2019
  • Wu Minhua +4
  • Research Article
  • Citations49

GFCC based discriminatively trained noise robust continuous ASR system for Hindi language

  • May 07, 2018
  • Journal of Ambient Intelligence and Humanized Computing
  • Mohit Dua +2
  • Conference Article
  • Citations3

The Effect of Different Optimization Techniques on End-to-End Turkish Speech Recognition Systems that use Connectionist Temporal Classification

  • Oct 01, 2018
  • 2018 2nd International Symposium on Multidisciplinary Studies and Innovative Technologies (ISMSIT)
  • Recep Sinan Arslan +1
  • Book Chapter
  • Citations7

Advances in STC Russian Spontaneous Speech Recognition System

  • Jan 01, 2016
  • Ivan Medennikov +1
  • Conference Article
  • Citations3

A Deep Learning Approach for Bangla Speech to Text Conversion

  • Aug 16, 2021
  • Raima Adhikary +3
  • Research Article
  • Citations15

Improving Deep Learning based Automatic Speech Recognition for Gujarati

  • Dec 13, 2021
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Deepang Raval +3
  • Conference Article
  • Citations3

Coping with disfluencies in spontaneous speech recognition

  • Oct 04, 2004
  • Frederik Stouten +1
  • Research Article
  • Citations36

Advancing Acoustic-to-Word CTC Model With Attention and Mixed-Units

  • Sep 04, 2019
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Amit Das +4
  • Conference Article
  • Citations372

Improved Training of End-to-end Attention Models for Speech Recognition

  • Sep 02, 2018
  • Albert Zeyer +3
  • Conference Article
  • Citations6

A New Corpus of Elderly Japanese Speech for Acoustic Modeling, and a Preliminary Investigation of Dialect-Dependent Speech Recognition

  • Oct 01, 2019
  • Meiko Fukuda +4
  • Research Article
  • Citations21

Japanese large-vocabulary continuous-speech recognition using a newspaper corpus and broadcast news

  • Jun 01, 1999
  • Speech Communication
  • Katsutoshi Ohtsuki +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.