• Home
  • Search
  • Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition
  • Open Access IconOpen Access
  • Cite Icon54
  • https://doi.org/10.1109/icassp.2019.8682977Copy DOI Icon

Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition

  • May 1, 2019
  • Wu Minhua +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Conventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech enhancement techniques do not always yield ASR accuracy improvement because the optimization criterion for speech enhancement is not directly relevant to the ASR objective. In this work, we develop new acoustic modeling techniques that optimize spatial filtering and long short-term memory (LSTM) layers from multi-channel (MC) input based on an ASR criterion directly. In contrast to conventional methods, we incorporate array processing knowledge into the acoustic model. Moreover, we initialize the network with beamformers' coefficients. We investigate effects of such MC neural networks through ASR experiments on the real-world far-field data where users are interacting with an ASR system in uncontrolled acoustic environments. We show that our MC acoustic model can reduce a word error rate (WER) by~16.5\% compared to a single channel ASR system with the traditional log-mel filter bank energy (LFBE) feature on average. Our result also shows that our network with the spatial filtering layer on two-channel input achieves a relative WER reduction of~9.5\% compared to conventional beamforming with seven microphones.

Similar Papers
  • Research Article
  • Citations25

Enhancements in automatic Kannada speech recognition system by background noise elimination and alternate acoustic modelling

  • Jan 22, 2020
  • International Journal of Speech Technology
  • G Thimmaraja Yadava +1
  • Research Article
  • Citations25

Combined speech enhancement and auditory modelling for robust distributed speech recognition

  • May 20, 2008
  • Speech Communication
  • Ronan Flynn +1
  • Book Chapter

Frame-Level Selective Decoding Using Native and Non-native Acoustic Models for Robust Speech Recognition to Native and Non-native Speech

  • Aug 28, 2013
  • Yoo Rhee Oh +3
  • Book Chapter
  • Citations30

Impact of Age in ASR for the Elderly: Preliminary Experiments in European Portuguese

  • Jan 01, 2012
  • Thomas Pellegrini +5
  • Conference Article
  • Citations21

Graph-based semi-supervised acoustic modeling in DNN-based speech recognition

  • Dec 01, 2014
  • Yuzong Liu +1
  • Conference Article
  • Citations36

Some insights from translating conversational telephone speech

  • May 01, 2014
  • Gaurav Kumar +3
  • Conference Article
  • Citations13

Improving Character Error Rate is Not Equal to Having Clean Speech: Speech Enhancement for ASR Systems with Black-Box Acoustic Models

  • May 23, 2022
  • Ryosuke Sawata +2
  • Research Article
  • Citations6

Croatian Large Vocabulary Automatic Speech Recognition

  • Jan 18, 2017
  • Automatika
  • Sanda Martinčić-Ipšić +2
  • Research Article
  • Citations15

Improving Deep Learning based Automatic Speech Recognition for Gujarati

  • Dec 13, 2021
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Deepang Raval +3
  • Conference Article
  • Citations41

Transliteration Based Approaches to Improve Code-Switched Speech Recognition Performance

  • Dec 01, 2018
  • Jesse Emond +4
  • Conference Article
  • Citations3

Using Taigi Dramas with Mandarin Chinese Subtitles to Improve Taigi Speech Recognition

  • Nov 05, 2020
  • Pin-Yuan Chen +5
  • Conference Article
  • Citations8

Can continuous speech recognizers handle isolated speech?

  • Sep 22, 1997
  • Fil Alleva +3
  • PDF
  • Research Article
  • Citations8

End-to-end automated speech recognition using a character based small scale transformer architecture

  • May 01, 2024
  • Expert Systems With Applications
  • Alexander Loubser +2
  • Research Article
  • Citations49

GFCC based discriminatively trained noise robust continuous ASR system for Hindi language

  • May 07, 2018
  • Journal of Ambient Intelligence and Humanized Computing
  • Mohit Dua +2
  • Research Article
  • Citations41

Multi-microphone speech recognition integrating beamforming, robust feature extraction, and advanced DNN/RNN backend

  • Feb 27, 2017
  • Computer Speech & Language
  • Takaaki Hori +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.