ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion
arXiv:1811.01609
Abstract
This paper proposes a voice conversion (VC) method using sequence-to-sequence (seq2seq or S2S) learning, which flexibly converts not only the voice characteristics but also the pitch contour and duration of input speech. The proposed method, called ConvS2S-VC, has three key features. First, it uses a model with a fully convolutional architecture. This is particularly advantageous in that it is suitable for parallel computations using GPUs. It is also beneficial since it enables effective normalization techniques such as batch normalization to be used for all the hidden layers in the networks. Second, it achieves many-to-many conversion by simultaneously learning mappings among multiple speakers using only a single model instead of separately learning mappings between each speaker pair using a different model. This enables the model to fully utilize available training data collected from multiple speakers by capturing common latent features that can be shared across different speakers. Owing to this structure, our model works reasonably well even without source speaker information, thus making it able to handle any-to-many conversion tasks. Third, we introduce a mechanism, called the conditional batch normalization that switches batch normalization layers in accordance with the target speaker. This particular mechanism has been found to be extremely effective for our many-to-many conversion model. We conducted speaker identity conversion experiments and found that ConvS2S-VC obtained higher sound quality and speaker similarity than baseline methods. We also found from audio examples that it could perform well in various tasks including emotional expression conversion, electrolaryngeal speech enhancement, and English accent conversion.
Published in IEEE/ACM Trans. ASLP https://ieeexplore.ieee.org/document/9113442
References in corpus (15)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Sequence to Sequence Learning with Neural Networks
- WaveNet: A Generative Model for Raw Audio
- Attention-Based Models for Speech Recognition
- A Learned Representation For Artistic Style
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention
- Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks
- Efficient Neural Audio Synthesis
- The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods
- ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech
- ACVAE-VC: Non-parallel many-to-many voice conversion with auxiliary classifier variational autoencoder
- StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks
- FloWaveNet : A Generative Flow for Raw Audio
- Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech Separation
Cited by in corpus (7)
- StarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice Conversion
- Measuring the Effectiveness of Voice Conversion on Speaker Identification and Automatic Speech Recognition Systems
- CycleGAN-VC2: Improved CycleGAN-based Non-parallel Voice Conversion
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- ConVoice: Real-Time Zero-Shot Voice Style Transfer with Convolutional Network
- Pretraining Techniques for Sequence-to-Sequence Voice Conversion
- Non-autoregressive sequence-to-sequence voice conversion