Voice Conversion from Unaligned Corpora using Variational Autoencoding Wasserstein Generative Adversarial Networks
arXiv:1704.00849
Abstract
Building a voice conversion (VC) system from non-parallel speech corpora is challenging but highly valuable in real application scenarios. In most situations, the source and the target speakers do not repeat the same texts or they may even speak different languages. In this case, one possible, although indirect, solution is to build a generative model for speech. Generative models focus on explaining the observations with latent variables instead of learning a pairwise transformation function, thereby bypassing the requirement of speech frame alignment. In this paper, we propose a non-parallel VC framework with a variational autoencoding Wasserstein generative adversarial network (VAW-GAN) that explicitly considers a VC objective when building the speech model. Experimental results corroborate the capability of our framework for building a VC system from unaligned data, and demonstrate improved conversion quality.
Submitted to INTERSPEECH 2017
Cited by in corpus (57)
- An Introduction to Variational Autoencoders
- How Generative Adversarial Networks and Their Variants Work: An Overview
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Deep Voice 2: Multi-Speaker Neural Text-to-Speech
- AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
- F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder
- Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling
- Cycle-Consistent Speech Enhancement
- Voice Conversion Based on Cross-Domain Features Using Variational Auto Encoders
- StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks
- Adversarial Feature-Mapping for Speech Enhancement
- Adversarial Training in Affective Computing and Sentiment Analysis: Recent Advances and Perspectives
- Deep learning methods in speaker recognition: a review
- An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
- Unpaired Speech Enhancement by Acoustic and Adversarial Supervision for Speech Recognition
- One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization
- Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data
- Towards adversarial learning of speaker-invariant representation for speech emotion recognition
- Adversarial representation learning for private speech generation
- Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder
- Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion
- VQVAE Unsupervised Unit Discovery and Multi-scale Code2Spec Inverter for Zerospeech Challenge 2019
- Seen and Unseen emotional style transfer for voice conversion with a new emotional speech dataset
- VAW-GAN for Singing Voice Conversion with Non-parallel Training Data
- Attentive Adversarial Learning for Domain-Invariant Training
- Many-to-Many Voice Transformer Network
- Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations
- Emotional Voice Conversion With Cycle-consistent Adversarial Network
- Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet
- Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training
- Emotional Voice Conversion: Theory, Databases and ESD
- Investigation of Using Disentangled and Interpretable Representations for One-shot Cross-lingual Voice Conversion
- Statistical Parametric Speech Synthesis Using Generative Adversarial Networks Under A Multi-task Learning Framework
- NoiseVC: Towards High Quality Zero-Shot Voice Conversion
- Generative Adversarial Networks for Unpaired Voice Transformation on Impaired Speech
- CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-spectrogram Conversion
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- AGAIN-VC: A One-shot Voice Conversion using Activation Guidance and Adaptive Instance Normalization
- VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-shot Voice Conversion
- Transfer Learning from Monolingual ASR to Transcription-free Cross-lingual Voice Conversion
- Synthetic Epileptic Brain Activities Using Generative Adversarial Networks
- Crossmodal Voice Conversion
- Building Multi lingual TTS using Cross Lingual Voice Conversion
- Error Reduction Network for DBLSTM-based Voice Conversion
- Training Generative Adversarial Networks with Adaptive Composite Gradient
- MASS: Multi-task Anthropomorphic Speech Synthesis Framework
- Many-to-Many Voice Conversion with Out-of-Dataset Speaker Support
- Expressive Voice Conversion: A Joint Framework for Speaker Identity and Emotional Style Transfer
- Transferring Source Style in Non-Parallel Voice Conversion
- Learn distributed GAN with Temporary Discriminators
- GAZEV: GAN-Based Zero-Shot Voice Conversion over Non-parallel Speech Corpus
- Simulating dysarthric speech for training data augmentation in clinical speech applications
- Rhythm-Flexible Voice Conversion without Parallel Data Using Cycle-GAN over Phoneme Posteriorgram Sequences
- Many-to-Many Voice Conversion using Cycle-Consistent Variational Autoencoder with Multiple Decoders
- Learning a Generative Model of Cancer Metastasis
- Learning in your voice: Non-parallel voice conversion based on speaker consistency loss
- Spectrum and Prosody Conversion for Cross-lingual Voice Conversion with CycleGAN