• Home
  • Search
  • Automatic Speaker Recognition by Speech Signal
  • Cite Icon5
  • https://doi.org/10.5772/6333Copy DOI Icon

Automatic Speaker Recognition by Speech Signal

  • Oct 1, 2008
  • Milan Sigmund
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Acoustical communication is one of the fundamental prerequisites for the existence of human society. Textual language has become extremely important in modern life, but speech has dimensions of richness that text cannot approximate. From speech alone, fairly accurate guesses can be made as to whether the speaker is male or female, adult or child. In addition, experts can extract from speech information regarding e.g. the speaker’s state of mind. As computer power increased and knowledge about speech signals improved, research of speech processing became aimed at automated systems for many purposes. Speaker recognition is the complement of speech recognition. Both techniques use similar methods of speech signal processing. In automatic speech recognition, the speech processing approach tries to extract linguistic information from the speech signal to the exclusion of personal information. Conversely, speaker recognition is focused on the characteristics unique to the individual, disregarding the current word spoken. The uniqueness of an individual’s voice is a consequence of both the physical features of the person vocal tract and the person mental ability to control the muscles in the vocal tract. An ideal speaker recognition system would use only physical features to characterize speakers, since these features cannot be easily changed. However, it is obvious that the physical features as vocal tract dimensions of an unknown speaker cannot be simply measured. Thus, numerical values for physical features or parameters would have to be derived from digital signal processing parameters extracted from the speech signal. Suppose that vocal tracts could be effectively represented by 10 independent physical features, with each feature taking on one of 10 discrete values. In this case, 1010 individuals in the population (i.e., 10 billion) could be distinguished whereas today’s world population amounts to approximately 7 billion individuals. People can reliably identify familiar voices. About 2-3 seconds of speech is sufficient to identify a voice, although performance decreases for unfamiliar voices. One review of human speaker recognition (Lancker et al., 1985) notes that many studies of 8-10 speakers (work colleagues) yield in excess of 97% accuracy if a sentence or more of the test speech is heard. Performance falls to about 54% when duration is shorter than 1 second and/or distorted e.g., severely highpass or lowpass filtered. Performance also falls significantly if training and test utterances are processed through different transmission systems. A study

Similar Papers
  • Book Chapter

An Overview of the Concept of Speaker Recognition

  • Dec 06, 2019
  • Intelligent Systems
  • Sindhu Rajendran +5
  • Research Article
  • Citations15

Human-computer interaction for virtual-real fusion

  • Jan 01, 2023
  • Journal of Image and Graphics
  • Jianhua Tao +5
  • Conference Article
  • Citations5

A realtime implementation of a text independent speaker recognition system

  • Apr 01, 1981
  • E Wrench
  • Research Article
  • Citations2

Simultaneous speaker identification and watermarking

  • Jan 15, 2021
  • International Journal of Speech Technology
  • Basant S Abd El-Wahab +3
  • Conference Article
  • Citations9

Speech recognition of different sampling rates using fractal code descriptor

  • Jul 01, 2016
  • Rattaphon Hokking +2
  • Research Article
  • Citations15

A study on model-based error rate estimation for automatic speech recognition

  • Nov 01, 2003
  • IEEE Transactions on Speech and Audio Processing
  • Chao-Shih Huang +2
  • PDF
  • Research Article
  • Citations3

An Impact of Narrowband Speech Codec Mismatch on a Performance of GMM-UBM Speaker Recognition over Telecommunication Channel

  • Feb 29, 2016
  • Communications - Scientific letters of the University of Zilina
  • Jozef Polacky +2
  • Research Article
  • Citations103

Fifty years of progress in speech and speaker recognition

  • Oct 01, 2004
  • The Journal of the Acoustical Society of America
  • Sadaoki Furui
  • Conference Article
  • Citations7

Mel frequency cepstral coefficients based text independent Automatic Speaker Recognition using matlab

  • Feb 01, 2014
  • Amit Kumar Singh +2
  • Book Chapter

Modelling and Understanding of Speech and Speaker Recognition

  • Sep 12, 2011
  • Tilendra Shishir +1
  • Conference Article
  • Citations4

AANN models for speaker recognition based on difference cepstrals

  • Jul 20, 2003
  • S Guruprasad +2
  • Dissertation

Noise Reduction with Microphone Arrays for Speaker Identification

  • Dec 01, 2012
  • Zachary Gideon Cohen
  • Research Article

Neural Network Architectures for Extracting Meaningful Representations from Audio Data

  • Nov 21, 2025
  • International Journal of Scientific Research in Engineering and Management
  • M Shubha +3
  • Book Chapter

Speaker Specific Formant Dynamics of Vowels

  • Jan 01, 2020
  • Sharada Vikram Chougule
  • Research Article
  • Citations7

Residual networks for text-independent speaker identification: Unleashing the power of residual learning

  • Dec 18, 2023
  • Journal of Information Security and Applications
  • Pooja Gambhir +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.