ConVoice: Real-Time Zero-Shot Voice Style Transfer with Convolutional Network
arXiv:2005.07815
Abstract
We propose a neural network for zero-shot voice conversion (VC) without any parallel or transcribed data. Our approach uses pre-trained models for automatic speech recognition (ASR) and speaker embedding, obtained from a speaker verification task. Our model is fully convolutional and non-autoregressive except for a small pre-trained recurrent neural network for speaker encoding. ConVoice can convert speech of any length without compromising quality due to its convolutional architecture. Our model has comparable quality to similar state-of-the-art models while being extremely fast.
References in corpus (6)
- AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
- F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder
- Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining
- QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions
- SqueezeWave: Extremely Lightweight Vocoders for On-device Speech Synthesis
- Non-Parallel Voice Conversion with Cyclic Variational Autoencoder