StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks
arXiv:1806.02169
Abstract
This paper proposes a method that allows non-parallel many-to-many voice conversion (VC) by using a variant of a generative adversarial network (GAN) called StarGAN. Our method, which we call StarGAN-VC, is noteworthy in that it (1) requires no parallel utterances, transcriptions, or time alignment procedures for speech generator training, (2) simultaneously learns many-to-many mappings across different attribute domains using a single generator network, (3) is able to generate converted speech signals quickly enough to allow real-time implementations and (4) requires only several minutes of training examples to generate reasonably realistic-sounding speech. Subjective evaluation experiments on a non-parallel many-to-many speaker identity conversion task revealed that the proposed method obtained higher sound quality and speaker similarity than a state-of-the-art method based on variational autoencoding GANs.
References in corpus (6)
- WaveNet: A Generative Model for Raw Audio
- Neural Discrete Representation Learning
- Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks
- The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods
- Voice Conversion from Non-parallel Corpora Using Variational Auto-encoder
- Generative adversarial network-based approach to signal reconstruction from magnitude spectrograms
Cited by in corpus (22)
- AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
- F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder
- Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
- ACVAE-VC: Non-parallel many-to-many voice conversion with auxiliary classifier variational autoencoder
- Unsupervised Speech Decomposition via Triple Information Bottleneck
- Semi-blind source separation with multichannel variational autoencoder
- An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
- WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions
- ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion
- One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization
- HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
- WaveCycleGAN2: Time-domain Neural Post-filter for Speech Waveform Generation
- Non-Parallel Voice Conversion with Cyclic Variational Autoencoder
- Anisotropic Stroke Control for Multiple Artists Style Transfer
- Measuring the Effectiveness of Voice Conversion on Speaker Identification and Automatic Speech Recognition Systems
- ConVoice: Real-Time Zero-Shot Voice Style Transfer with Convolutional Network
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- AttS2S-VC: Sequence-to-Sequence Voice Conversion with Attention and Context Preservation Mechanisms
- FastVC: Fast Voice Conversion with non-parallel data
- Many-to-Many Voice Conversion with Out-of-Dataset Speaker Support
- Multivariate-Information Adversarial Ensemble for Scalable Joint Distribution Matching
- Rhythm-Flexible Voice Conversion without Parallel Data Using Cycle-GAN over Phoneme Posteriorgram Sequences