• Home
  • Search
  • Collaborative Joint Training With Multitask Recurrent Model for Speech and Speaker Recognition
  • Cite Icon66
  • https://doi.org/10.1109/taslp.2016.2639323Copy DOI Icon

Collaborative Joint Training With Multitask Recurrent Model for Speech and Speaker Recognition

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Automatic speech and speaker recognition are traditionally treated as two independent tasks and are studied separately. The human brain in contrast deciphers the linguistic content, and the speaker traits from the speech in a collaborative manner. This key observation motivates the work presented in this paper. A collaborative joint training approach based on multitask recurrent neural network models is proposed, where the output of one task is backpropagated to the other tasks. This is a general framework for learning collaborative tasks and fits well with the goal of joint learning of automatic speech and speaker recognition. Through a comprehensive study, it is shown that the multitask recurrent neural net models deliver improved performance on both automatic speech and speaker recognition tasks as compared to single-task systems. The strength of such multitask collaborative learning is analyzed, and the impact of various training configurations is investigated.

Similar Papers
  • Research Article

Design and Development of an Advanced Speaker Recognition System Using MFCC and Neural Network

  • Feb 16, 2026
  • International Journal of Innovative Research in Computer and Communication Engineering
  • Pawan Kamble +2
  • Conference Article
  • Citations6

Comparing audio and visual information for speech processing

  • Aug 28, 2005
  • D Dean +3
  • Research Article
  • Citations103

Fifty years of progress in speech and speaker recognition

  • Oct 01, 2004
  • The Journal of the Acoustical Society of America
  • Sadaoki Furui
  • Research Article
  • Citations25

Combined speech enhancement and auditory modelling for robust distributed speech recognition

  • May 20, 2008
  • Speech Communication
  • Ronan Flynn +1
  • Research Article
  • Citations24

Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition

  • Jun 26, 2023
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Qingyu Wang +5
  • Conference Article
  • Citations2

Train Your Classifier First: Cascade Neural Networks Training from Upper Layers to Lower Layers

  • Jun 06, 2021
  • Shucong Zhang +5
  • PDF
  • Research Article
  • Citations3

A sub-band-based feature reconstruction approach for robust speaker recognition

  • Oct 21, 2014
  • EURASIP Journal on Audio, Speech, and Music Processing
  • Furong Yan +2
  • Book Chapter

An Overview of the Concept of Speaker Recognition

  • Dec 06, 2019
  • Intelligent Systems
  • Sindhu Rajendran +5
  • Conference Article
  • Citations9

The modified group delay feature: a new spectral representation of speech

  • Oct 04, 2004
  • Hema A Murthy +2
  • Conference Article
  • Citations20

Emerging features for speaker recognition

  • Jan 01, 2007
  • Eliathamby Ambikairajah
  • Conference Article
  • Citations10

Empirical comparison of analog and digital auditory preprocessing for automatic speech recognition

  • Aug 07, 2002
  • T.M Massengill +3
  • Conference Article
  • Citations21

Graph-based semi-supervised acoustic modeling in DNN-based speech recognition

  • Dec 01, 2014
  • Yuzong Liu +1
  • Research Article
  • Citations3

CNN-Based Speaker Verification and Speech Recognition in Tibetan

  • Dec 01, 2020
  • Journal of Physics: Conference Series
  • Zhenye Gan +3
  • Conference Article
  • Citations1

End-to-end Oriental Language Speech Recognition with Integrated Language Identification

  • Oct 01, 2022
  • Anbin Qi +4
  • Book Chapter
  • Citations1

Lipreading Using Recurrent Neural Prediction Model

  • Jan 01, 2004
  • Takuya Tsunekawa +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.