• Home
  • Search
  • Multilevel sampling and aggregation for discriminative training
  • https://doi.org/10.1109/iscslp.2014.6936677Copy DOI Icon

Multilevel sampling and aggregation for discriminative training

  • Sep 1, 2014
  • Yunxin Zhao +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

We propose to use data sampling in the extended Baum-Welch (EBW) algorithm for maximum mutual information (MMI) based estimation of speech acoustic models, and to randomize the configurations of the sampled training sets and aggregate the numerator and denominator sufficient statistics for improving model robustness. We further combine data sampling based ensemble acoustic modeling with the data sampling based EBW, forming a two-level data sampling mechanism for acoustic model training. We conducted experiments on a telehealth conversational speech recognition task, where the two-level data sampling mechanism gave a statistically significant, absolute word accuracy gain of 3.56% over the conventional MMI baseline, corresponding to a 19.44% relative word error rate reduction.

Similar Papers
  • Conference Article
  • Citations135

Towards Fast and Accurate Streaming End-To-End ASR

  • May 01, 2020
  • Bo Li +6
  • Conference Article
  • Citations28

L-Vector: Neural Label Embedding for Domain Adaptation

  • Apr 10, 2020
  • Zhong Meng +6
  • Conference Article
  • Citations73

Semi-supervised GMM and DNN acoustic model training with multi-system combination and confidence re-calibration

  • Aug 25, 2013
  • Yan Huang +3
  • Conference Article
  • Citations3

Effective acoustic modeling for rate-of-speech variation in large vocabulary conversational speech recognition

  • Oct 04, 2004
  • Jing Zheng +2
  • Book Chapter
  • Citations4

A New Perspective on Combining GMM and DNN Frameworks for Speaker Adaptation

  • Jan 01, 2016
  • Natalia Tomashenko +2
  • Conference Article
  • Citations6

Exploring Retraining-free Speech Recognition for Intra-sentential Code-switching

  • May 01, 2019
  • Zhen Huang +5
  • Conference Article

Batch Normalization based Unsupervised Speaker Adaptation for Acoustic Models

  • Nov 01, 2019
  • Jiangyan Yi +1
  • Conference Article
  • Citations9

FMPE-MAP: improved discriminative adaptation for modeling new domains

  • Aug 27, 2007
  • Jing Zheng +1
  • Research Article
  • Citations49

GFCC based discriminatively trained noise robust continuous ASR system for Hindi language

  • May 07, 2018
  • Journal of Ambient Intelligence and Humanized Computing
  • Mohit Dua +2
  • Conference Article
  • Citations33

How Does Pre-Trained Wav2Vec 2.0 Perform on Domain-Shifted Asr? an Extensive Benchmark on Air Traffic Control Communications

  • Jan 09, 2023
  • Juan Zuluaga-Gomez +8
  • Book Chapter
  • Citations8

Exploring GMM-derived Features for Unsupervised Adaptation of Deep Neural Network Acoustic Models

  • Jan 01, 2016
  • Natalia Tomashenko +3
  • Conference Article
  • Citations23

Comparison of optimization methods for discriminative training criteria

  • Sep 22, 1997
  • Ralf Schlüter +4
  • Conference Article
  • Citations7

Exploring Machine Speech Chain For Domain Adaptation

  • May 23, 2022
  • Fengpeng Yue +4
  • Conference Article
  • Citations5

Multi-Channel Opus Compression for Far-Field Automatic Speech Recognition with a Fixed Bitrate Budget

  • Aug 30, 2021
  • Lukas Drude +3
  • Conference Article
  • Citations135

LSTM time and frequency recurrence for automatic speech recognition

  • Dec 01, 2015
  • Jinyu Li +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.