• Home
  • Search
  • A Parallel-Data-Free Speech Enhancement Method Using Multi-Objective Learning Cycle-Consistent Generative Adversarial Network
  • Cite Icon52
  • https://doi.org/10.1109/taslp.2020.2997118Copy DOI Icon

A Parallel-Data-Free Speech Enhancement Method Using Multi-Objective Learning Cycle-Consistent Generative Adversarial Network

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Recently, deep neural networks (DNNs) have become the mainstream strategy for speech enhancement task because it can achieve the higher speech quality and intelligibility than the traditional methods. However, these DNN-based methods always need a large number of parallel corpus consisting of clean speech and noise to produce noisy data for the training of the DNN in order to improve the generalization of the network. As a result, this implies that many noisy speech signals that are collected in real environment cannot be used to train the DNN because of the lack of corresponding clean speech and noise. Additionally, as we know, noise varies with the time and scenario, so we cannot obtain parallel speech and noise due to infinite noise data and some limited speech data. Thus, the network training with unparallel speech and noise data is essential for the generalization of the network. To address this problem, we propose a novel parallel-data-free speech enhancement method, in which the cycle-consistent generative adversarial network (CycleGAN) and multi-objective learning are employed. Our method is also able to make best use of the benefits of multi-objective learning. On the training stage, we utilize two different encoders to encode the features of clean speech and noisy speech, respectively. Then, two forward generators are immediately used to predict the ideal time-frequency (T-F) mask and log-power spectrum (LPS) of clean speech. Two inverse generators are applied to map the magnitude spectrum (MS) and LPS of noisy speech, respectively. In addition, four discriminators are used to distinguish the real speech features from the generated features. Two encoders, four generators and four discriminators are simultaneously trained by using adversarial, identity-mapping, latent similarity and cycle-consistent loss. On the test stage, we directly utilize the forward generators and encoders to acquire the enhanced speech. The experimental results indicate that the proposed approach is able to achieve the better speech enhancement performance than the reference methods. Moreover, the proposed method is also effective to improve speech quality and intelligibility when the networks are trained under the parallel data.

Similar Papers
  • PDF
  • Research Article
  • Citations20

Speech enhancement from fused features based on deep neural network and gated recurrent unit network

  • Oct 24, 2021
  • EURASIP Journal on Advances in Signal Processing
  • Youming Wang +3
  • Conference Article
  • Citations7

Speech enhancement based on noise compensated magnitude spectrum

  • May 01, 2014
  • Md T Islam +4
  • Research Article

Objective and Subjective Evaluation of Speech Enhancement Methods in the UDASE task of the 7th CHIME Challenge

  • May 31, 2025
  • INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
  • Gongalla Mayuri
  • Supplementary Content
  • Citations1

Analysis of very low quality speech for mask-based enhancement

  • Sep 01, 2013
  • Spiral (Imperial College London)
  • Sira González
  • Conference Article

Monaural Speech Enhancement using Deep Neural Network with Cross-Speech Dataset

  • Sep 13, 2021
  • Norezmi Jamal +3
  • Book Chapter

A Speech Enhancement Method Combining Two-Branch Communication and Spectral Subtraction

  • Jan 01, 2023
  • Ruhan He +4
  • Research Article
  • Citations2

Dynamic controllable speech enhancement models based on quantile loss functions

  • Nov 09, 2023
  • Applied Acoustics
  • Wenhao Yuan +3
  • Conference Article
  • Citations8

DNN-Based Speech Enhancement Using MBE Model

  • Sep 01, 2018
  • Qizheng Huang +3
  • Research Article
  • Citations21

A Joint Framework of Denoising Autoencoder and Generative Vocoder for Monaural Speech Enhancement

  • Jan 01, 2020
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Zhihao Du +2
  • Research Article
  • Citations10

Spectral Phase Estimation Based on Deep Neural Networks for Single Channel Speech Enhancement

  • Dec 01, 2019
  • Journal of Communications Technology and Electronics
  • N Saleem +2
  • PDF
  • Research Article
  • Citations16

Improved Relativistic Cycle-Consistent GAN With Dilated Residual Network and Multi-Attention for Speech Enhancement

  • Jan 01, 2020
  • IEEE Access
  • Yutian Wang +4
  • Conference Article
  • Citations102

Exploring multi-channel features for denoising-autoencoder-based speech enhancement

  • Apr 01, 2015
  • Shoko Araki +5
  • Conference Article
  • Citations1

Variance normalized perceptual subspace speech enhancement with noise estimation using SPP

  • Sep 01, 2016
  • Sudeep Surendran +1
  • Conference Article
  • Citations1

HMM-based cue parameters estimation for speech enhancement

  • Oct 01, 2016
  • Feng Deng +2
  • Conference Article
  • Citations10

Single channel speech enhancement utilizing iterative processing of multi-band spectral subtraction algorithm

  • Dec 01, 2012
  • Navneet Upadhyay +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.