- Research Article
575
- 10.1016/s0165-1684(01)00128-1
Speech enhancement for non-stationary noise environments
- Oct 05, 2001
- Signal Processing
- Israel Cohen + 1 more +1
Speech enhancement for non-stationary noise environments
In this study, a novel human acoustic perception motivated Wiener filter speech enhancement system is presented to cope with real-world interfering background noises. Guiding by the speech presence probability, two alternative methods are proposed to reduce the noise by adopting the audible sound pressure level (SPL) and the masking characteristic of the human auditory system to achieve better listening comfort level. More specifically, when the probability of speech presence in the noisy signal is less than the decision threshold, a new SPL compressed method effectively reduces the noise. When the speech presence probability is more than the decision threshold, an improved acoustical mask threshold constrained Wiener filter approach enhances the noisy speech. Moreover, in order to evaluate the performance of the new system, the proposed algorithm is compared with the classic prior signal-to-noise ratio-based Wiener filter and three acoustic perception related algorithms. The experimental results show that the proposed algorithm significantly outperforms the four comparing algorithms in terms of speech quality and intelligibility either in stationary or moderate non-stationary noisy environments. Thus, the intended approach can be employed as the front-end module for various speech-related applications.
Speech enhancement for non-stationary noise environments
Speech enhancement for non-stationary noise environments
Variance normalized perceptual subspace speech enhancement with noise estimation using SPP
In this paper, a perceptual subspace speech enhancement method using masking properties of human auditory system is proposed. Noise was estimated at every frame using Speech Presence Probability (SPP). Variance normalization was further done to remove the abrupt changes in the noisy speech signal, removing noise. A combined masking threshold is calculated from simultaneous and temporal masking, which is used to decide the gain parameters for the algorithm. Spectral Domain Constrained (SDC) estimator was employed in determining the filter coefficients. The performance evaluation was done using wcep and WSS objective measures and subjective informal listening test. The results show a superior performance of the proposed method over some of the existing speech enhancement methods.
Read moreSpeech enhancement for nonstationary noise environments
In this paper, we propose a robust speech enhancement algorithm for nonstationary noise environments, which comprises a multiband spectral subtraction (MBSS) speech estimator and a minima controlled recursive averaging (MCRA) noise estimation. A multiband approach is obtained by dividing the whole spectrum into multiple subbands and applying spectral subtraction independently in each band. The noise estimate exploit the observation that the noise signal typically has a nonuniform effect on the spectrum of speech and is given by averaging the past spectrum power value and using a time and frequency dependent smoothing factor that is calculated based on the speech presence probability in subbands. We test the performance of proposed algorithm in various nonstationary noise environments. The experiment results confirm the superiority of the MBSS and MCRA estimator. Outstanding speech enhancement is achieved, while avoiding the residual musical noise phenomena.
Read moreDual Microphone Speech Enhancement Based on Statistical Modeling of Interchannel Phase Difference
The interchannel phase difference (IPD) may be one of the most widely-used spatial cues in multichannel speech processing, and has been used in beamformers and post filters for speech enhancement. The coherence, which is also used as a feature for speech enhancement, can provide information on the reliability of the IPD for the estimation of the speech presence probability (SPP). In this paper, we propose dual microphone speech enhancement adopting <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">a posteriori</i> SPP estimation based on statistical modeling of the IPD. The marginal distribution of the IPD is derived from the distribution of the relative transfer function which is parameterized with the IPD and coherence, with a single assumption that the observed discrete Fourier transform (DFT) coefficients in each frequency are distributed according to a complex bivariate Gaussian distribution. Given the direction of arrival of the desired signal, the <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">a posteriori</i> SPP is obtained using the IPD distributions with and without the information on the location of the interfering source, and is applied to speech enhancement. Experimental results for various types and locations of noise, signal-to-noise ratios, reverberation times, and locations of the target source showed that the proposed method outperformed previously proposed approaches utilizing IPD information.
Read moreHybrid probabilistic adaptation mode controller for generalized sidelobe cancellers applied to multi-microphone speech enhancement
Hybrid probabilistic adaptation mode controller for generalized sidelobe cancellers applied to multi-microphone speech enhancement
Read moreImproved speech presence probability estimation based on wavelet denoising
A reliable estimator for speech presence probability (SPP) can significantly improve the performance of many speech enhancement algorithms. Previous work showed that a good SPP estimator can be obtained by using a smooth a-posteriori signal to noise ratio (SNR) function, which can be achieved by reducing the noise variance when estimating the speech power spectrum. In this paper, a wavelet based denoising algorithm is proposed for such purpose. We first apply the wavelet transform to the periodogram of a noisy speech signal to generate an oracle for indicating the locations of the noise floor in the periodogram. We then make use of that oracle to selectively remove the wavelet coefficients of the noise floor in the log multitaper spectrum (MTS) of the noisy speech. The remaining wavelet coefficients are then used to reconstruct a denoised MTS and in turn generate a smooth a-posteriori SNR function. Simulation results show that the new SPP estimator outperforms the traditional approaches and enables a significantly improvement in the quality and intelligibility of the enhanced speeches. © 2012 IEEE.
Read moreData-Driven Non-Intrusive Speech Intelligibility Prediction Using Speech Presence Probability
Time consuming Speech Intelligibility (SI) listening tests with human subjects can be replaced by algorithmic SI predictors. In recent years, data-driven SI predictors have been showing promising results. A major limiting factor in the advancement of data-driven SI prediction is that there is a scarcity of SI listening test data available to train the data-driven methods. In this article we propose a data-driven SI predictor that does not require access to an underlying noise-free reference signal, i.e., <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">non-intrusive</i> , and which does not require listening test data for training. Instead, the proposed method exploits a hypothesized link between SI and Speech Presence Probability (SPP). We show that a neural network can be trained on easily obtainable speech in additive noise data to estimate SPP, and that a simple post-processing stage can be applied in order to map the estimated SPP to SI predictions with high accuracy. The proposed method is evaluated and compared to other state-of-the art non-intrusive SI predictors, and achieves the highest performance even in the presence of processed noisy speech, which the SPP estimator has not been trained on.
Read moreSpeech enhancement using fourth-order cumulants and optimum filters in the subband domain
Speech enhancement using fourth-order cumulants and optimum filters in the subband domain
Speech enhancement via two-stage dual tree complex wavelet packet transform with a speech presence probability estimator.
In this paper, a two-stage dual tree complex wavelet packet transform (DTCWPT) based speech enhancement algorithm has been proposed, in which a speech presence probability (SPP) estimator and a generalized minimum mean squared error (MMSE) estimator are developed. To overcome the drawback of signal distortions caused by down sampling of wavelet packet transform (WPT), a two-stage analytic decomposition concatenating undecimated wavelet packet transform (UWPT) and decimated WPT is employed. An SPP estimator in the DTCWPT domain is derived based on a generalized Gamma distribution of speech, and Gaussian noise assumption. The validation results show that the proposed algorithm can obtain enhanced perceptual evaluation of speech quality (PESQ), and segmental signal-to-noise ratio (SegSNR) at low signal-to-noise ratio (SNR) nonstationary noise, compared with four other state-of-the-art speech enhancement algorithms, including optimally modified log-spectral amplitude (OM-LSA), soft masking using a posteriori SNR uncertainty (SMPO), a posteriori SPP based MMSE estimation (MMSE-SPP), and adaptive Bayesian wavelet thresholding (BWT).
Read moreUsing the turbo principle for exploiting temporal and spectral correlations in speech presence probability estimation
In this paper we present a speech presence probability (SPP) estimation algorithmwhich exploits both temporal and spectral correlations of speech. To this end, the SPP estimation is formulated as the posterior probability estimation of the states of a two-dimensional (2D) Hidden Markov Model (HMM). We derive an iterative algorithm to decode the 2D-HMM which is based on the turbo principle. The experimental results show that indeed the SPP estimates improve from iteration to iteration, and further clearly outperform another state-of-the-art SPP estimation algorithm.
Read moreDnn-Based Ar-Wiener Filtering for Speech Enhancement
This paper presents a novel approach for estimating autoregressive (AR) model parameters using deep neural network (DNN) in the AR-Wiener filtering speech enhancement. Unlike conventional DNN that predicts one kind of target, the DNN used in this paper is trained to predict the AR model parameters of speech and noise simultaneously at offline stage. We train this network by minimizing the Euclidean distance between the output of DNN and the AR model parameters of clean speech and noise. At online stage, the acoustic features are first extracted from noisy speech as the input of the DNN. Then, AR model parameters of speech and noise are estimated by the DNN simultaneously. Finally, the Wiener filter is constructed by the AR model parameters of speech and noise. However, the AR model parameters only models the spectral shape not the spectral details, there are still some residual noise between the harmonics. In order to solve this problem, we introduce the speech-presence probability (SPP), that is, in the test stage, the SPP is estimated and is used to update the Wiener filter. The experimental results show that our approach has higher performance compared with some existing approaches.
Read moreEstimation of speech absence uncertainty based on multiple linear regression analysis for speech enhancement
Estimation of speech absence uncertainty based on multiple linear regression analysis for speech enhancement
Dual-Microphone Speech Dereverberation in a Noisy Environment
Speech signals recorded with a distant microphone usually contain reverberation and noise, which degrade the fidelity and intelligibility of speech, and the recognition performance of automatic speech recognition systems. In E. Habets (2005) presented a multi-microphone speech dereverberation algorithm to suppress late reverberation in a noise-free environment. In this paper we show how an estimate of the late reverberant energy can be obtained from noisy observations. A more sophisticated speech enhancement technique based on the optimally-modified log spectral amplitude (OM-LSA) estimator is used to suppress the undesired late reverberant signal and noise. The speech presence probability used in the OM-LSA is extended to improve the decision between speech, late reverberation and noise. Experiments using simulated and real acoustic impulse responses are presented and show significant reverberation reduction with little speech distortion
Read moreDNN-based speech mask estimation for eigenvector beamforming
In this paper, we present an optimal multi-channel Wiener filter, which consists of an eigenvector beamformer and a single-channel postfilter. We show that both components solely depend on a speech presence probability, which we learn using a deep neural network, consisting of a deep autoencoder and a softmax regression layer. To prevent the DNN from learning specific speaker and noise types, we do not use the signal energy as input feature, but rather the cosine distance between the dominant eigenvectors of consecutive frames of the power spectral density of the noisy speech signal. We compare our system against the BeamformIt toolkit, and state-of-the-art approaches such as the front-end of the best system of the CHiME3 challenge. We show that our system yields superior results, both in terms of perceptual speech quality and classification error.
Read moreA Speech Enhancement System for Automotive Speech Recognition with a Hybrid Voice Activity Detection Method
This paper presents a front-end speech enhancement approach to robust speech recognition in automotive environments. It combines hybrid voice activity detection (VAD), relative transfer function (RT-F) based generalized sidelobe cancelation, and single-channel post filtering to enhance the speech signal of interest, thereby improving the robustness of speech recognition. First, we choose four typical driving scenarios, which include most of the noise types in automobiles to record training data. The recorded data is then used to train deep neural network models (DNNs) for both speech and noise. The trained DNNs are subsequently used to estimate the speech presence probability on a frame-by-frame basis. This speech presence probability is then combined with the output of an energy-based VAD to form a hybrid VAD, which serves as the basis for the rest components of the speech enhancement system, including RTF estimation, adaptive beamforming, and post-filtering. Experiments are conducted in real automotive environments. The results show that the developed method can significantly improve the performance of both VAD and automatic speech recognition (ASR).
Read more