The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods
arXiv:1804.04262
Abstract
We present the Voice Conversion Challenge 2018, designed as a follow up to the 2016 edition with the aim of providing a common framework for evaluating and comparing different state-of-the-art voice conversion (VC) systems. The objective of the challenge was to perform speaker conversion (i.e. transform the vocal identity) of a source speaker to a target speaker while maintaining linguistic information. As an update to the previous challenge, we considered both parallel and non-parallel data to form the Hub and Spoke tasks, respectively. A total of 23 teams from around the world submitted their systems, 11 of them additionally participated in the optional Spoke task. A large-scale crowdsourced perceptual evaluation was then carried out to rate the submitted converted speech in terms of naturalness and similarity to the target speaker identity. In this paper, we present a brief summary of the state-of-the-art techniques for VC, followed by a detailed explanation of the challenge tasks and the results that were obtained.
Accepted for Speaker Odyssey 2018
Cited by in corpus (22)
- A Survey on Speech Deepfake Detection
- ACVAE-VC: Non-parallel many-to-many voice conversion with auxiliary classifier variational autoencoder
- StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks
- Semi-blind source separation with multichannel variational autoencoder
- Unsupervised Representation Disentanglement using Cross Domain Features and Adversarial Learning in Variational Autoencoder based Voice Conversion
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching References
- Quasi-Periodic Parallel WaveGAN: A Non-autoregressive Raw Waveform Generative Model with Pitch-dependent Dilated Convolution Neural Network
- ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion
- Preech: A System for Privacy-Preserving Speech Transcription
- DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake Voices
- Non-Parallel Voice Conversion with Cyclic Variational Autoencoder
- Many-to-Many Voice Transformer Network
- Emotional Voice Conversion: Theory, Databases and ESD
- FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion
- SpeechNet: A Universal Modularized Model for Speech Processing Tasks
- ConVoice: Real-Time Zero-Shot Voice Style Transfer with Convolutional Network
- Collapsed speech segment detection and suppression for WaveNet vocoder
- DiDiSpeech: A Large Scale Mandarin Speech Corpus
- Nonparallel Voice Conversion with Augmented Classifier Star Generative Adversarial Networks
- Expressive Voice Conversion: A Joint Framework for Speaker Identity and Emotional Style Transfer
- How do Voices from Past Speech Synthesis Challenges Compare Today?
- Low-Latency Real-Time Non-Parallel Voice Conversion based on Cyclic Variational Autoencoder and Multiband WaveRNN with Data-Driven Linear Prediction