• Home
  • Search
  • Robust Feature Extraction Using Modulation Filtering of Autoregressive Models
  • Cite Icon49
  • https://doi.org/10.1109/taslp.2014.2329190Copy DOI Icon

Robust Feature Extraction Using Modulation Filtering of Autoregressive Models

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Speaker and language recognition in noisy and degraded channel conditions continue to be a challenging problem mainly due to the mismatch between clean training and noisy test conditions. In the presence of noise, the most reliable portions of the signal are the high energy regions which can be used for robust feature extraction. In this paper, we propose a front end processing scheme based on autoregressive (AR) models that represent the high energy regions with good accuracy followed by a modulation filtering process. The AR model of the spectrogram is derived using two separable time and frequency AR transforms. The first AR model (temporal AR model) of the sub-band Hilbert envelopes is derived using frequency domain linear prediction (FDLP). This is followed by a spectral AR model applied on the FDLP envelopes. The output 2-D AR model represents a low-pass modulation filtered spectrogram of the speech signal. The band-pass modulation filtered spectrograms can further be derived by dividing two AR models with different model orders (cut-off frequencies). The modulation filtered spectrograms are converted to cepstral coefficients and are used for a speaker recognition task in noisy and reverberant conditions. Various speaker recognition experiments are performed with clean and noisy versions of the NIST-2010 speaker recognition evaluation (SRE) database using the state-of-the-art speaker recognition system. In these experiments, the proposed front-end analysis provides substantial improvements (relative improvements of up to 25%) compared to baseline techniques. Furthermore, we also illustrate the generalizability of the proposed methods using language identification (LID) experiments on highly degraded high-frequency (HF) radio channels and speech recognition experiments on noisy data.

Similar Papers
  • Conference Article
  • Citations6

Robust speaker recognition using spectro-temporal autoregressive models

  • Aug 25, 2013
  • Sri Harish Mallidi +2
  • Conference Article
  • Citations2

Frequency Domain Linear Prediction-based robust text-dependent speaker identification

  • Oct 01, 2016
  • M A Islam
  • Research Article
  • Citations46

Significance of analytic phase of speech signals in speaker verification

  • Feb 26, 2016
  • Speech Communication
  • Karthika Vijayan +2
  • Research Article
  • Citations49

Local spectral variability features for speaker verification

  • Nov 18, 2015
  • Digital Signal Processing
  • Md Sahidullah +1
  • Research Article
  • Citations4

Relevance factor of maximum a posteriori adaptation for GMM–NAP–SVM in speaker and language recognition

  • Sep 29, 2014
  • Computer Speech & Language
  • Chang Huai You +2
  • Conference Article
  • Citations7

Frame selection of interview channel for NIST speaker recognition evaluation

  • Nov 01, 2010
  • Hanwu Sun +2
  • Conference Article
  • Citations25

CNN based speaker recognition in language and text-independent small scale system

  • Dec 01, 2019
  • Rohan Jagiasi +3
  • Conference Article
  • Citations22

Softsad: Integrated frame-based speech confidence for speaker recognition

  • Apr 01, 2015
  • Mitchell Mclaren +2
  • Conference Article
  • Citations4

Language Identification in Overlapped Multi-lingual Speeches

  • Jul 22, 2022
  • Zuhragvl Aysa +2
  • Conference Article
  • Citations34

The QUT-NOISE-SRE protocol for the evaluation of noisy speaker recognition

  • Sep 06, 2015
  • David Dean +4
  • Conference Article
  • Citations27

The MIT-LL/IBM 2006 Speaker Recognition System: High-Performance Reduced-Complexity Recognition

  • Apr 01, 2007
  • W M Campbell +4
  • Conference Article
  • Citations39

Mean hilbert envelope coefficients (MHEC) for robust speaker recognition

  • Sep 09, 2012
  • Seyed Omid Sadjadi +2
  • Conference Article
  • Citations6

Comparison of modulation features for phoneme recognition

  • Jan 01, 2010
  • Sriram Ganapathy +2
  • PDF
  • Research Article
  • Citations3

A sub-band-based feature reconstruction approach for robust speaker recognition

  • Oct 21, 2014
  • EURASIP Journal on Audio, Speech, and Music Processing
  • Furong Yan +2
  • Research Article
  • Citations13

Investigation of the effect of data duration and speaker gender on text-independent speaker recognition

  • Oct 27, 2012
  • Computers & Electrical Engineering
  • Cemal Hanilçi +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.