• Home
  • Search
  • Learning Speech Representations with Flexible Hidden Feature Dimensions
  • Cite Icon4
  • https://doi.org/10.1109/icassp49357.2023.10094969Copy DOI Icon

Learning Speech Representations with Flexible Hidden Feature Dimensions

  • Jun 4, 2023
  • Huaizhen Tang +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Non-parallel many-to-many voice conversion is a kind of style transfer task in speech. Recently, AutoVC has been applied in this field as a popular solution, as it can achieve distribution-matching style transfer by training only the re- construction loss. However, in order to strike a good balance between timbre disentanglement and sound quality, AutoVC requires imposing very strict constraints on the dimensionality of the latent representation. This constraint affects the quality of the converted speech while making it challenging to apply to other datasets directly. This paper proposes a new voice conversion framework that uses only one encoder to obtain timbre and content information by partitioning the latent space in the channel dimension. Furthermore, two different types of classifiers and two additional reconstruction losses are proposed to ensure that different parts of the latent space contain only separated content and timbre information, respectively. Experiments on the VCTK dataset show that the proposed model achieves state-of-the-art results in terms of the naturalness and similarity of converted speech. In addition, we experimentally show that for different division proportions of latent space, the content and timbre information will always be well separated.

Similar Papers
  • Conference Article
  • Citations7

Voicy: Zero-Shot Non-Parallel Voice Conversion in Noisy Reverberant Environments

  • Aug 26, 2021
  • Alejandro Mottini +3
  • Research Article
  • Citations72

User identity linkage across social networks via linked heterogeneous network embedding

  • Apr 23, 2018
  • World Wide Web
  • Yaqing Wang +5
  • Conference Article

Manipulating Image Style Transformation via Latent-Space SVM

  • Oct 01, 2021
  • Qiudan Wang
  • Conference Article
  • Citations2

A New Spoken Language Teaching Tech: Combining Multi-attention and AdaIN for One-shot Cross Language Voice Conversion

  • Dec 11, 2022
  • Dengfeng Ke +5
  • Research Article
  • Citations6

Diffusion‐based Human Motion Style Transfer with Semantic Guidance

  • Oct 09, 2024
  • Computer Graphics Forum
  • Lei Hu +4
  • Conference Article
  • Citations84

One-Shot Voice Conversion by Vector Quantization

  • May 01, 2020
  • Da-Yi Wu +1
  • Conference Article
  • Citations16

Artistic Style Discovery with Independent Components

  • Jun 01, 2022
  • Xin Xie +5
  • Conference Article
  • Citations25

Disentangling Content and Fine-Grained Prosody Information Via Hybrid ASR Bottleneck Features for Voice Conversion

  • May 23, 2022
  • Xintao Zhao +6
  • Research Article
  • Citations7

Learning Many-to-Many Mapping for Unpaired Real-World Image Super-Resolution and Downscaling.

  • Dec 01, 2024
  • IEEE transactions on pattern analysis and machine intelligence
  • Wanjie Sun +1
  • Research Article

Improving 4D Seismic History Matching Through Data Analysis: A Localized Sensitivity Analysis Workflow

  • Jun 20, 2024
  • SPE Journal
  • Rasool A Kolajoobi +2
  • Conference Article
  • Citations2

Speak Like a Dog: Human to Non-human creature Voice Conversion

  • Nov 07, 2022
  • Kohei Suzuki +3
  • Research Article
  • Citations30

Play as You Like: Timbre-Enhanced Multi-Modal Music Style Transfer

  • Jul 17, 2019
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Chien-Yu Lu +4
  • Conference Article
  • Citations3

Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice Conversion

  • Oct 26, 2023
  • Yanzhen Ren +5
  • Research Article
  • Citations30

Interlayer Link Prediction in Multiplex Social Networks Based on Multiple Types of Consistency Between Embedding Vectors.

  • Apr 01, 2023
  • IEEE Transactions on Cybernetics
  • Rui Tang +5
  • Conference Article
  • Citations7

Reconstructing Dual Learning for Neural Voice Conversion Using Relatively Few Samples

  • Dec 13, 2021
  • Aolan Sun +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.