• Home
  • Search
  • Non-Parallel Training in Voice Conversion Using an Adaptive Restricted Boltzmann Machine
  • Cite Icon79
  • https://doi.org/10.1109/taslp.2016.2593263Copy DOI Icon

Non-Parallel Training in Voice Conversion Using an Adaptive Restricted Boltzmann Machine

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In this paper, we present a voice conversion (VC) method that does not use any parallel data while training the model. VC is a technique where only speaker-specific information in source speech is converted while keeping the phonological information unchanged. Most of the existing VC methods rely on parallel data—pairs of speech data from the source and target speakers uttering the same sentences. However, the use of parallel data in training causes several problems: 1) the data used for the training are limited to the predefined sentences, 2) the trained model is only applied to the speaker pair used in the training, and 3) mismatches in alignment may occur. Although it is, thus, fairly preferable in VC not to use parallel data, a nonparallel approach is considered difficult to learn. In our approach, we achieve nonparallel training based on a speaker adaptation technique and capturing latent phonological information. This approach assumes that speech signals are produced from a restricted Boltzmann machine-based probabilistic model, where phonological information and speaker-related information are defined explicitly. Speaker-independent and speaker-dependent parameters are simultaneously trained under speaker adaptive training. In the conversion stage, a given speech signal is decomposed into phonological and speaker-related information, the speaker-related information is replaced with that of the desired speaker, and then voice-converted speech is obtained by mixing the two. Our experimental results showed that our approach outperformed another nonparallel approach, and produced results similar to those of the popular conventional Gaussian mixture models-based method that used parallel data in subjective and objective criteria.

Similar Papers
  • Conference Article
  • Citations15

A dual alignment scheme for improved speech-to-singing voice conversion

  • Dec 01, 2017
  • Karthika Vijayan +2
  • Research Article
  • Citations1

Speech naturalness improvement via $$\mathrm {\epsilon }$$ ϵ -closed extended vectors sets in voice conversion systems

  • Jan 12, 2017
  • Multidimensional Systems and Signal Processing
  • Mohammad Javad Jannati +2
  • Research Article
  • Citations8

Voice conversion with SI-DNN and KL divergence based mapping without parallel training data

  • Nov 30, 2018
  • Speech Communication
  • Feng-Long Xie +2
  • Conference Article
  • Citations5

End-to-End Voice Conversion with Information Perturbation

  • Dec 11, 2022
  • Qicong Xie +4
  • Conference Article
  • Citations3

Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice Conversion

  • Oct 26, 2023
  • Yanzhen Ren +5
  • Conference Article
  • Citations11

Two-Stage Training Method for Japanese Electrolaryngeal Speech Enhancement Based on Sequence-to-Sequence Voice Conversion

  • Jan 09, 2023
  • Ding Ma +3
  • Conference Article
  • Citations6

A Survey on Generative Adversarial Networks based Models for Many-to-many Non-parallel Voice Conversion

  • Mar 09, 2022
  • Yasmin Alaa +2
  • Conference Article
  • Citations4

Statistical acoustic-to-articulatory mapping unified with speaker normalization based on voice conversion

  • Sep 06, 2015
  • Hidetsugu Uchida +3
  • Book Chapter
  • Citations3

First Steps Towards New Czech Voice Conversion System

  • Jan 01, 2006
  • Zdeněk Hanzlíček +1
  • Conference Article
  • Citations7

Non-Parallel Many-To-Many Voice Conversion by Knowledge Transfer from a Text-To-Speech Model

  • Jun 06, 2021
  • Xinyuan Yu +1
  • Research Article
  • Citations13

Confidence Score Based Speaker Adaptation of Conformer Speech Recognition Systems

  • Jan 01, 2023
  • IEEE/ACM Transactions on Audio, Speech, and Language Processing
  • Jiajun Deng +8
  • Conference Article
  • Citations14

Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification Systems

  • Jun 04, 2023
  • Danwei Cai +2
  • Book Chapter
  • Citations4

Voice Conversion by Mapping the Spectral and Prosodic Features Using Support Vector Machine

  • Jan 01, 2009
  • Rabul Hussain Laskar +3
  • Conference Article
  • Citations1

A GMM based residual prediction method for voice conversion

  • Jan 01, 2005
  • Jing Xia +1
  • Research Article
  • Citations2

Pretraining and Fine-Tuning Techniques for Electrolaryngeal Speech Enhancement Based on Sequence-to-Sequence Voice Conversion

  • Jan 01, 2025
  • IEEE Transactions on Audio, Speech and Language Processing
  • Ding Ma +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.