• Home
  • Search
  • State-Clustering Based Multiple Deep Neural Networks Modeling Approach for Speech Recognition
  • Cite Icon31
  • https://doi.org/10.1109/taslp.2015.2392944Copy DOI Icon

State-Clustering Based Multiple Deep Neural Networks Modeling Approach for Speech Recognition

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The hybrid deep neural network (DNN) and hidden Markov model (HMM) has recently achieved dramatic performance gains in automatic speech recognition (ASR). The DNN-based acoustic model is very powerful but its learning process is extremely time-consuming. In this paper, we propose a novel DNN-based acoustic modeling framework for speech recognition, where the posterior probabilities of HMM states are computed from multiple DNNs (mDNN), instead of a single large DNN, for the purpose of parallel training towards faster turnaround. In the proposed mDNN method all tied HMM states are first grouped into several disjoint clusters based on data-driven methods. Next, several hierarchically structured DNNs are trained separately in parallel for these clusters using multiple computing units (e.g. GPUs). In decoding, the posterior probabilities of HMM states can be calculated by combining outputs from multiple DNNs. In this work, we have shown that the training procedure of the mDNN under popular criteria, including both frame-level cross-entropy and sequence-level discriminative training, can be parallelized efficiently to yield significant speedup. The training speedup is mainly attributed to the fact that multiple DNNs are parallelized over multiple GPUs and each DNN is smaller in size and trained by only a subset of training data. We have evaluated the proposed mDNN method on a 64-hour Mandarin transcription task and the 320-hour Switchboard task. Compared to the conventional DNN, a 4-cluster mDNN model with similar size can yield comparable recognition performance in Switchboard (only about 2% performance degradation) with a greater than 7 times speed improvement in CE training and a 2.9 times improvement in sequence training, when 4 GPUs are used.

Similar Papers
  • PDF
  • Research Article
  • Citations5

Integrating Expression Data-Based Deep Neural Network Models with Biological Networks to Identify Regulatory Modules for Lung Adenocarcinoma.

  • Aug 30, 2022
  • Biology
  • Lei Fu +12
  • Conference Article
  • Citations39

Joint Optimization of DNN Partition and Scheduling for Mobile Cloud Computing

  • Aug 09, 2021
  • Yubin Duan +1
  • Book Chapter
  • Citations25

Deep Neural Network Based Continuous Speech Recognition for Serbian Using the Kaldi Toolkit

  • Jan 01, 2015
  • Branislav Popović +4
  • Conference Article
  • Citations13

Distance-aware DNNs for robust speech recognition

  • Sep 06, 2015
  • Yajie Miao +1
  • Conference Article
  • Citations36

On decomposing a deep neural network into modules

  • Nov 07, 2020
  • Rangeet Pan +1
  • Research Article
  • Citations5

Dialect Identification in Telugu Language Speech Utterance Using Modified Features with Deep Neural Network

  • Dec 31, 2021
  • Traitement du Signal
  • Shivaprasad Satla +1
  • Research Article
  • Citations20

Optimizing makespan and resource utilization for multi-DNN training in GPU cluster

  • Jun 24, 2021
  • Future Generation Computer Systems
  • Zhongjin Li +5
  • Research Article
  • Citations58

Real-time model predictive cooling control for an HVAC system in a factory building

  • Feb 09, 2023
  • Energy and Buildings
  • Seon Jung Ra +2
  • Research Article
  • Citations53

HiTDL: High-Throughput Deep Learning Inference at the Hybrid Mobile Edge

  • Dec 01, 2022
  • IEEE Transactions on Parallel and Distributed Systems
  • Jing Wu +5
  • Research Article
  • Citations17

FTBME: feature transferring based multi-model ensemble

  • Mar 12, 2020
  • Multimedia Tools and Applications
  • A Yongquan Yang +4
  • PDF
  • Research Article
  • Citations10

An Adaptive Task Migration Scheduling Approach for Edge‐Cloud Collaborative Inference

  • Jan 01, 2022
  • Wireless Communications and Mobile Computing
  • Boyin Zhang +4
  • Conference Article
  • Citations2

Syllable-based Speech Recognition for a Very Low-Resource Language, Chaha

  • Dec 20, 2019
  • Tessfu Geteye Fantaye +2
  • Research Article
  • Citations185

An Empirical Study of the Impact of Hyperparameter Tuning and Model Optimization on the Performance Properties of Deep Neural Networks

  • Apr 09, 2022
  • ACM Transactions on Software Engineering and Methodology
  • Lizhi Liao +3
  • Book Chapter
  • Citations1

Handwritten Digit Recognition Using Very Deep Convolutional Neural Network

  • Jan 01, 2022
  • M Dhilsath Fathima +2
  • Research Article
  • Citations5

SSAT: Active Authorization Control and User’s Fingerprint Tracking Framework for DNN IP Protection

  • Oct 29, 2024
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Mingfu Xue +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.