• Home
  • Search
  • Automatic Speaker Verification on Compressed Audio
  • https://doi.org/10.1109/dessert58054.2022.10018734Copy DOI Icon

Automatic Speaker Verification on Compressed Audio

  • Dec 9, 2022
  • Oleksandra Sokol +5 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Voice-based phishing and wire-fraud attacks have become a topical problem in the recent years due to the emergence of advanced AI-based speech synthesis models. These models can generate realistic speech signal of a known target that is difficult to differentiate from a bonafide voice of real human. This proved to be an issue by the recent security reports related to vishing attacks on bank call centers, fraud or pranking of public figures, and spoofing of voice authentication systems. Current approaches to address voice fraud issue are based on applying an Automatic Speaker Verification (ASV) system. In most cases, these systems are tuned on datasets that consist of wideband quality bonafide and spoofed voice samples. This makes ASV systems vulnerable to speech signal degradation caused by voice encoding in cellular network and Voice over IP (VoIP). However, performance evaluation of ASV (namely Equal Error Rate (EER) estimation) is almost exclusively available only for the cellular networks. Thus, performance of ASV systems for modern VoIP applications remains unclear. In this paper, we evaluate the modern ASV systems on audio compressed with codecs used in both cellular networks (AMR and GSM codecs) and VoIP applications (G.711, G.722, AAC, Lyra and Opus codecs). In addition, ASV performance was tested using popular VoIP application (Discord). Obtained results have shown that codec application results in considerable (up to two times) EER increase compared to the baseline results. Moreover, we observed up to three times increase in EER on data transmitted using Discord. We propose to apply hard samples mining to the training process in order to improve the accuracy of ASV systems on compressed voice samples. It allows to reduce EER from 21% down to 16% even for the most distorted samples obtained after aggressive voice compression by GSM codecs. Note, that improvement for real VoIP application is even higher - with EER on Discord data decrease from 35% to 20%.

Similar Papers
  • Research Article
  • Citations16

Children's speaker verification in low and zero resource conditions

  • Jun 07, 2021
  • Digital Signal Processing
  • S Shahnawazuddin +3
  • Research Article
  • Citations37

Deep Learning Serves Voice Cloning: How Vulnerable Are Automatic Speaker Verification Systems to Spoofing Trials?

  • Feb 01, 2020
  • IEEE Communications Magazine
  • Pavol Partila +4
  • Research Article

Over-the-Air Adversarial Attacks and Detection for Automatic Speaker Verification

  • Jan 01, 2026
  • IEEE Transactions on Audio, Speech and Language Processing
  • Li Wang +5
  • Conference Article
  • Citations13

Multi-task learning of deep neural networks for joint automatic speaker verification and spoofing detection

  • Nov 01, 2019
  • Jiakang Li +2
  • Research Article
  • Citations111

Spoofing Detection in Automatic Speaker Verification Systems Using DNN Classifiers and Dynamic Acoustic Features.

  • Dec 04, 2017
  • IEEE Transactions on Neural Networks and Learning Systems
  • Hong Yu +4
  • Conference Article
  • Citations2

Automatic Speaker Verification System Substantiating Children’s Dialects in School Settings

  • Nov 25, 2022
  • Virender Kadyan +3
  • Research Article
  • Citations2

Enhancing Voice Authentication with a Hybrid Deep Learning and Active Learning Approach for Deepfake Detection

  • Nov 08, 2024
  • Journal of Robotics and Control (JRC)
  • Ali Saadoon Ahmed +1
  • Research Article
  • Citations23

Deep generative variational autoencoding for replay spoof detection in automatic speaker verification

  • Mar 19, 2020
  • Computer Speech & Language
  • Bhusan Chettri +2
  • Conference Article
  • Citations5

Automatic Speaker Verification and Replay Attack Detection System using novel Glottal Flow Cepstrum Coefficients

  • Dec 01, 2021
  • Yusra Banaras +2
  • Conference Article
  • Citations39

Introducing i-vectors for joint anti-spoofing and speaker verification

  • Sep 14, 2014
  • Elie Khoury +4
  • Research Article
  • Citations31

Detection of replay spoof speech using teager energy feature cues

  • Aug 14, 2020
  • Computer Speech & Language
  • Madhu R Kamble +1
  • Book Chapter
  • Citations2

Investigating Language Variability on the Performance of Speaker Verification Systems

  • Jan 01, 2018
  • Amir Vaheb +3
  • Research Article
  • Citations7

Detection of replay signals using excitation source and shifted CQCC features

  • Feb 04, 2021
  • International Journal of Speech Technology
  • Krishna Dutta +2
  • Conference Article
  • Citations205

Deep Residual Neural Networks for Audio Spoofing Detection

  • Sep 15, 2019
  • Moustafa Alzantot +2
  • Conference Article
  • Citations33

Voice Spoofing Countermeasure for Synthetic Speech Detection

  • Apr 05, 2021
  • Farman Hassan +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.