• Home
  • Search
  • Automatic speech recognition with an adaptation model motivated by auditory processing
  • Cite Icon70
  • https://doi.org/10.1109/tsa.2005.860349Copy DOI Icon

Automatic speech recognition with an adaptation model motivated by auditory processing

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The mel-frequency cepstral coefficient (MFCC) or perceptual linear prediction (PLP) feature extraction typically used for automatic speech recognition (ASR) employ several principles which have known counterparts in the cochlea and auditory nerve: frequency decomposition, mel- or bark-warping of the frequency axis, and compression of amplitudes. It seems natural to ask if one can profitably employ a counterpart of the next physiological processing step, synaptic adaptation. We, therefore, incorporated a simplified model of short-term adaptation into MFCC feature extraction. We evaluated the resulting ASR performance on the AURORA 2 and AURORA 3 tasks, in comparison to ordinary MFCCs, MFCCs processed by RASTA, and MFCCs processed by cepstral mean subtraction (CMS), and both in comparison to and in combination with Wiener filtering. The results suggest that our approach offers a simple, causal robustness strategy which is competitive with RASTA, CMS, and Wiener filtering and performs well in combination with Wiener filtering. Compared to the structurally related RASTA, our adaptation model provides superior performance on AURORA 2 and, if Wiener filtering is used prior to both approaches, on AURORA 3 as well.

Similar Papers
  • Conference Article
  • Citations300

A novel approach for MFCC feature extraction

  • Dec 01, 2010
  • Md Afzal Hossan +2
  • Conference Article
  • Citations8

A framework for robust MFCC feature extraction using SNR-dependent compression of enhanced mel filter bank energies

  • Sep 17, 2006
  • Babak Nasersharif +1
  • Research Article
  • Citations2

Research on the UAV Sound Recognition Method Based on Frequency Band Feature Extraction

  • May 05, 2025
  • Drones
  • Jilong Zhong +4
  • Research Article
  • Citations2

Frequency domain analysis of MFCC feature extraction in children’s speech recognition system

  • Feb 26, 2022
  • JURNAL INFOTEL
  • Risanuri Hidayat
  • Conference Article
  • Citations59

Denoising Speech for MFCC Feature Extraction Using Wavelet Transformation in Speech Recognition System

  • Jul 01, 2018
  • Risanuri Hidayat +3
  • Research Article
  • Citations49

GFCC based discriminatively trained noise robust continuous ASR system for Hindi language

  • May 07, 2018
  • Journal of Ambient Intelligence and Humanized Computing
  • Mohit Dua +2
  • Research Article
  • Citations46

Significance of analytic phase of speech signals in speaker verification

  • Feb 26, 2016
  • Speech Communication
  • Karthika Vijayan +2
  • Conference Article
  • Citations5

Efficient MFCC feature extraction on graphics processing units

  • Jan 01, 2013
  • Haofeng Kou +3
  • Book Chapter
  • Citations4

Gujarati Language Automatic Speech Recognition Using Integrated Feature Extraction and Hybrid Acoustic Model

  • Jan 01, 2023
  • Mohit Dua +1
  • PDF
  • Research Article
  • Citations102

Text-Independent Speaker Identification Through Feature Fusion and Deep Neural Network

  • Jan 01, 2020
  • IEEE Access
  • Rashid Jahangir +7
  • Research Article
  • Citations47

Application of a semi-automated vocal fingerprinting approach to monitor Bornean gibbon females in an experimentally fragmented landscape in Sabah, Malaysia

  • Jan 16, 2018
  • Bioacoustics
  • Dena J Clink +2
  • Conference Article
  • Citations1

Comparison between normalizations for SVM — GMM supervectors speaker verification

  • Nov 01, 2010
  • Udi Ben Simon +2
  • Research Article
  • Citations275

Fusion of Heterogeneous Speaker Recognition Systems in the STBU Submission for the NIST Speaker Recognition Evaluation 2006

  • Sep 01, 2007
  • IEEE Transactions on Audio, Speech, and Language Processing
  • Niko Brummer +9
  • Conference Article
  • Citations7

Particle Swarm Optimisation of Mel-frequency Cepstral Coefficients computation for the classification of asphyxiated infant cry

  • Oct 01, 2010
  • A Zabidi +4
  • Research Article
  • Citations1

Music Genre Classification Using Mel Frequency Cepstral Coefficients and Artificial Neural Networks: A Novel Approach

  • Dec 16, 2024
  • Scientific Journal of Informatics
  • Alamsyah Alamsyah +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.