- Research Article
31
- 10.1016/j.csl.2020.101140
Detection of replay spoof speech using teager energy feature cues
- Aug 14, 2020
- Computer Speech & Language
- Madhu R Kamble + 1 more +1
Detection of replay spoof speech using teager energy feature cues
In this paper, we explore the use of the deep learning approach for replay spoof detection in speaker verification systems. Automatic speaker verifications (ASVs) can be easily spoofed by previously recorded genuine speech. In order to counter the issues of spoofing, detecting spoofing attacks play an important role. Hence, we consider the detection of replay attack spoofing that is the most easily accomplished spoofing attack. In this light, we propose a deep neural network-based (DNN) classifier using a hybrid feature from Mel-frequency cepstral coefficient (MFCC) and constant Q cepstral coefficient (CQCC). Several experiments were conducted on the latest version of ASVspoof 2017 dataset. The results are compared with a base line system that uses the Gaussian mixture model (GMM) classifier with different features that include MFCC, CQCC, and the hybrid feature of the two. The experiment results reveal that the DNN classifier outperforms the conventional GMM classifier. It was found that the hybrid-based features are superior to single features, such as CQCC and MFCC in terms of equal error rate (ERR). In addition, like many previous researchers have found, it turned out that high-frequency regions of speech utterance convey much more discriminative information for replay attack detection.
Detection of replay spoof speech using teager energy feature cues
Detection of replay spoof speech using teager energy feature cues
Spoofing Detection in Automatic Speaker Verification Systems Using DNN Classifiers and Dynamic Acoustic Features.
With the development of speech synthesis technology, automatic speaker verification (ASV) systems have encountered the serious challenge of spoofing attacks. In order to improve the security of ASV systems, many antispoofing countermeasures have been developed. In the front-end domain, much research has been conducted on finding effective features which can distinguish spoofed speech from genuine speech and the published results show that dynamic acoustic features work more effectively than static ones. In the back-end domain, Gaussian mixture model (GMM) and deep neural networks (DNNs) are the two most popular types of classifiers used for spoofing detection. The log-likelihood ratios (LLRs) generated by the difference of human and spoofing log-likelihoods are used as spoofing detection scores. In this paper, we train a five-layer DNN spoofing detection classifier using dynamic acoustic features and propose a novel, simple scoring method only using human log-likelihoods (HLLs) for spoofing detection. We mathematically prove that the new HLL scoring method is more suitable for the spoofing detection task than the classical LLR scoring method, especially when the spoofing speech is very similar to the human speech. We extensively investigate the performance of five different dynamic filter bank-based cepstral features and constant Q cepstral coefficients (CQCC) in conjunction with the DNN-HLL method. The experimental results show that, compared to the GMM-LLR method, the DNN-HLL method is able to significantly improve the spoofing detection accuracy. Compared with the CQCC-based GMM-LLR baseline, the proposed DNN-HLL model reduces the average equal error rate of all attack types to 0.045%, thus exceeding the performance of previously published approaches for the ASVspoof 2015 Challenge task. Fusing the CQCC-based DNN-HLL spoofing detection system with ASV systems, the false acceptance rate on spoofing attacks can be reduced significantly.
Read moreSubband Analysis for Performance Improvement of Replay Attack Detection in Speaker Verification Systems
Automatic speaker verification systems have been widely employed in a variety of commercial applications. However, advancements in the field of speech technology have equipped the attackers with sophisticated techniques for circumventing speaker verification systems. The state-of-the-art countermeasures are fairly successful in detecting speech synthesis and voice conversion attacks. However, the problem of replay attack detection has not received much attention from the researchers. In this study, we perform subband analysis on constant-Q cepstral coefficient (CQCC) and mel-frequency cepstral coefficient (MFCC) features to improve the performance of replay attack detection. We have performed experiments on the ASVspoof 2017 database which consists of 3566 genuine and 15380 replay utterances. Our experimental results suggest that the features extracted from the high frequency band carries significant discriminatory information for replay attack detection. In particular, our approach achieves an improvement of 36.33% over the baseline replay attack detection method in terms of equal error rate.
Read moreApplying GMM-UBM framework for Indonesian forensic speaker verification
A text-dependent speaker verification system has been implemented for forensic application since 2008 when speech recording is allowed as legal evidence in Indonesian court. Due to the laborious procedures of the current system, an automatic text-independent speaker verification system is being developed in order to accelerate the forensic speaker verification process. The new forensic system is developed using a database spoken in Indonesian. The database is built on two scenarios, interview and conversation. This is done to simulate the real forensic condition for speaker verification in Indonesia. The system adopts the very famous Gaussian Mixture Model (GMM) and Universal Background Model (UBM) framework in speaker verification field and using Mel-Frequency Cepstral Coefficients (MFCC) as the speech feature. Along with the application of zero-normalization, the new forensic speaker verification system has successfully reached an Equal Error Rate (EER) of 4.66 % that is a better performance than the previous developed system of Indonesian forensic speaker verification.
Read moreDeep generative variational autoencoding for replay spoof detection in automatic speaker verification
Deep generative variational autoencoding for replay spoof detection in automatic speaker verification
Modelling Glottal Flow Derivative Signal for Detection of Replay Speech Samples
It is a widely known fact that automatic speaker verification systems are quite vulnerable to replay speech. The present work deals with detecting replay speech by using the information available in glottal flow derivative (GFD) signal. In signal processing terms, the speech signal can be represented as the response of a vocal-tract system with excited by a excitation source in the form of glottal flow. The effect of record and replay devices distorted the spectral characteristics of the naturally uttered speech sample, resulting distortion in corresponding GFD signals. In this work the GFD signals are parameterized by using standard mel filters and Gaussian mixtures models are made for detection. Although various methods are available, by correlation analysis it is observed that in the context of the present work the dynamic programming phase slope algorithm (DYPSA) method is relatively more effective in estimating the GFD signals. The experimental studies are made on ASVSpoof2017 database. The proposed glottal flow derivative mel frequency cepstral coefficients (GFDMFCC) feature provides 20.53% equal error rate (EER). This performance is comparatively poor than by speech and residual based features. It is mainly due to the absence of fine structure information in estimated GFD signal. However, in fusion with speech signal based constant-Q cepstral coefficients (CQCC) features, the GFDMFCC feature provides an improvement of 10.30% with reference to conventional residual feature. This shows the usefulness of modelling GFD signals for detection of replay signals.
Read moreComparison between normalizations for SVM — GMM supervectors speaker verification
This paper presents a comparison between several features normalization methods, and a comparison between different types of Gaussian Mixture Model (GMM) based supervectors normalizations for robust Speaker Verification. We implemented the methods of normalizations as a part of speaker verification system using Support Vector Machine (SVM) classifier and GMM-based supervectors. When implementing the speaker recognition system, we used Mel Frequency Cepstral Coefficients (MFCC) feature extraction. A valid question is which features normalization to use, if any. We examine the most common methods of feature normalizations, such as: Feature Warping mapping, and Cepstral Mean Subtraction (CMS) normalization with and without variance normalization. These methods were compared to features without normalization at all, and to a basic [-1, 1] normalization. In addition, we applied few types of normalizations to the GMM-mean supervectors, in order to improve the performance of the SVM classifier. All comparisons of the speaker verification system had been done in terms of DET curve, EER (Equal Error Rate) and Min. DCF. The best results we achieved were on combined supervector normalizations of Universal Background Model (UBM) Standard Deviation (STD) and [-1, 1] normalization. The type of the MFCC normalization has no big influence on the verification performance. The best results were: EER about 5.0% and MIN. DCF of 0.02.
Read moreDetection of replay signals using excitation source and shifted CQCC features
The replay attack is refereed as an unauthorized attempt to access the automatic speaker verification (ASV) system by using the pre-recorded speech samples of any target. The replay attack is performed by placing the pre-recorded speech sample of the target before the machine. Of late the replay attack is identified as the greatest threat to ASV system, mainly due to the availability of high quality recording and playback devices. In this work, excitation source feature referred as glottal mel frequency cepstral coefficient (GMFCC) and shifted constant Q cepstral coefficient (SCQCC) are proposed for detection of replay signals. The GMFCC is derived by applying conventional mel-cepstral technique to glottal flow derivative signal. The SCQCC is computed by using constant Q cepstral processing. The effectiveness of the proposed features are demonstrated by conducting experiments with ASVspoof 2017 version 2.0 database. The proposed GMFCC feature provides an equal error rate (EER) of 16.78%, that is 19.63% higher than the recently proposed residual mel frequency cepstral coefficient(RMFCC) feature. The conventional CQCC feature provides an EER of 12.32%. The proposed SCQCC feature provides an EER of 11.34%, shows a relative improvement of 7.94% over CQCC. Further, the CQCC in together with proposed GMFCC provides an EER of 8.82%. On the other hand, the proposed SCQCC+GMFCC system provides an EER of 8.60%. These results signify the usefulness of the proposed system to counter replay attacks.
Read moreA novel approach for MFCC feature extraction
The Mel-Frequency Cepstral Coefficients (MFCC) feature extraction method is a leading approach for speech feature extraction and current research aims to identify performance enhancements. One of the recent MFCC implementations is the Delta-Delta MFCC, which improves speaker verification. In this paper, a new MFCC feature extraction method based on distributed Discrete Cosine Transform (DCT-II) is presented. Speaker verification tests are proposed based on three different feature extraction methods including: conventional MFCC, Delta-Delta MFCC and distributed DCT-II based Delta-Delta MFCC with a Gaussian Mixture Model (GMM) classifier.
Read moreMixture linear prediction Gammatone Cepstral features for robust speaker verification under transmission channel noise
In this paper, we present a Mixture Linear Prediction based approach for robust Gammatone Cepstral Coefficients extraction (MLPGCCs). The proposed method provides performance improvement of Automatic Speaker Verification (ASV) using i-vector and Gaussian Probabilistic Linear Discriminant Analysis GPLDA modeling under transmission channel noise. The performance of the extracted MLPGCCs was evaluated using the NIST 2008 database where a single channel microphone recorded conversational speech. The system is analyzed in the presence of different channel transmission noises such as Additive White Gaussian (AWGN) and Rayleigh fading at various Signals to Noise Ratio (SNR) levels. The evaluation results show that the MLPGCCs features are a promising way for the ASV task. Indeed, the speaker verification performance using the MLPGCCs proposed features is significantly improved compared to the conventional Gammatone Frequency Cepstral Coefficients (GFCCs) and Mel Frequency Cepstral Coefficients (MFCCs) features. For speech signals corrupted with AWGN noise at SNRs ranging from (-5 dB to 15 dB), we obtain a significant reduction of the Equal Error Rate (EER) ranging from 9.41% to 6.65% and 3.72% to 1.50%, compared with conventional MFCCs and GFCCs features respectively. In addition, when the test speech signals are corrupted with Rayleigh fading channel we achieve an EER reduction ranging from 23.63% to 7.8% and from 10.88% to 6.8% compared with conventional MFCCs and GFCCs, respectively. We also found that the combination of GFCCs and MLPGCCs gives the highest performance of speaker verification system. The best performance combination achieved is around EER from 0.43% to 0.59% and 1.92% to 3.88%.
Read moreProcessing linear prediction residual signal to counter replay attacks
Replay attack is a method of using targets pre-recorded speech samples for acquiring unauthorized access to the automatic speaker verification (SV) system. This work investigates the usefulness of processing linear prediction (LP) residual signal to counter replay attacks. In detecting replay signals the major clues lie in tracing the record and playback devices characteristics that dominantly reflect at low frequency regions due to loud speaker, and at high frequency regions due to two-stage A/D conversions. Unlike speech, the spectral pattern of the impulse-like linear prediction (LP) residual signal is spread across entire frequency range, and so conjectured to be more useful. Based on the distribution nature of the mel-scale that tightly spaced in low frequency regions and reverse in inverse mel-scale, residual mel-frequency cepstral coefficients (RMFCC) and residual inverse mel-frequency cepstral coefficients (RIMFCC) features are used as the representative features. The effectiveness of these features is demonstrated on ASVspoof2017 database. The initial study is made on deciding the suitable prediction order. In terms of equal error rate (EER), RMFCC features provide the best performance of 14.57% from 20th order LP analysis and RIMFCC of 15.35% from 10th order LP analysis. These results show that residual signals effectively capture the low frequency distortions with higher order LP analysis and vice versa. The higher performance in case of RMFCC features may due to the significant distortion from loud speakers. The fusion of RMFCC and RIMFCC features further improves the performance to 10.14%, that is comparatively better than the state-of-the-art spectral centroid magnitude coefficients (SCMC) feature performance of 11.49%. Finally, the fusion of RMFCC and RIMFCC features together with SCMC provides 9.54%. These outcomes demonstrate the usefulness of processing LP residual signals to counter replay attacks.
Read moreGujarati Language Automatic Speech Recognition Using Integrated Feature Extraction and Hybrid Acoustic Model
In the case of low resource language, there is still the requirement for developing more efficient Automatic Speech Recognition (ASR) systems. In the proposed work, the ASR system is developed for the Gujarati language publicly available dataset. The approach in this paper applies the combination of Mel-frequency Cepstral Coefficients (MFCC) with Constant Q Cepstral Coefficients (CQCC)-based integrated front-end feature extraction techniques. To implement the backend part of the system, hybrid acoustic model is applied. Two-dimensional Convolutional Neural Network (Conv2D) with Bi-directional Gated Recurrent Units-based (BiGRU) backend model is used as the model. To build the ASR system, Connectionist Temporal Classification (CTC) loss function, CTC and prefix-based greedy decoder are also used with the acoustic model. The proposed work shows that the joint MFCC and CQCC feature extraction techniques show the 10–19% improvement in Word Error Rate (WER) as compared to isolated delta-delta features with the available integrated model.
Read moreSpeaker verification with multi-run ICA based speech enhancement
Forensic speaker verification systems show severe performance degradation in the presence of noise when the signal to noise ratio (SNR) is low. A possible solution to this problem is the use of multi-run independent component analysis (ICA) to reduce the effect of noise from the noisy speech signals. Previous works have used multi-run ICA in biosignal application; however, the effectiveness of multi-run ICA on noisy speaker verification has not been investigated yet. In this paper, we use multi-run ICA to enhance the noisy speech signals by choosing the highest signal to interference ratio (SIR) of the mixing matrix from different mixing matrices generated by iterating the fast ICA algorithm for several times. We use a combination of feature-warped mel frequency cepstral coefficients (MFCCs) and feature-warped MFCC extracted from the discrete wavelet transform (DWT) of the enhanced speech signals as the feature extraction. A state-of-the-art identity vector (i-vector) probabilistic linear discriminant analysis (PLDA) was used as a classifier in this paper. Experimental results demonstrate that the proposed method with multi-run ICA achieves high improvements in equal error rate (EER) of 66.68%, 69.24% and 70.78% over the baseline noisy speaker verification system, when the test speech signals are corrupted with CAR, STREET, and HOME noises respectively at −10 dB SNR.
Read moreSpeaker Verification Using IMNMF and MFCC with Feature Warping Under Noisy Environment
The mismatch between the training conditions and the test conditions severely degrades the performance of speaker verification. Aiming at solving this problem, this paper presents a method which uses a combination of improved nonnegative matrix factorization (IMNMF) and mel frequency cepstral coefficients (MFCC) with feature warping for improving identity-vector (i-vector) speaker verification performance in noisy environment. Unlike the traditional nonnegative matrix factorization (NMF), IMNMF uses extra free basis vectors to capture the features which are not included in training data, and linear constraints on dictionary atoms. Feature warping is used to remove channel noises. Therefore, the proposed method can reduce distortion of reconstructed speech while enhancing the recovered speech quality. The performance of i-vector speaker verification is evaluated by using the short utterance database and the NOISEX-92 database. The experiment results indicate that the score level fusion of feature-warped IMNMF-MFCC and feature-warped MFCC is superior to the baseline system at the equal error rate under the majority of signal-to-noise ratios (SNRs).
Read moreBimodal biometric person authentication using speech and face under degraded condition
In this work, we present a bimodal biometric system using speech and face features and tested its performance under degraded condition. Speaker verification (SV) system is built using Mel-Frequency Cepstral Coefficients (MFCC) followed by delta and delta-delta for feature extraction and Gaussian Mixture Model (GMM) for modeling. A face verification (FV) system is built using the combination of Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA). Sum rule is used for the fusion of the biometric scores. The performance of SV system under degraded condition is also checked. All the experimental results are shown upon a subset of IITG-DIT M4 multi-biometric database. The complementary information derived from the speech biometric at training stage is used to further decrease the FV error rate, which is termed as Cohort fed FV system. Finally we propose an improved bimodal person authentication system using SV and Cohort fed FV biometric systems.
Read more