Sample Efficient Adaptive Text-to-Speech
arXiv:1809.10460
Abstract
We present a meta-learning approach for adaptive text-to-speech (TTS) with few data. During training, we learn a multi-speaker model using a shared conditional WaveNet core and independent learned embeddings for each speaker. The aim of training is not to produce a neural network with fixed weights, which is then deployed as a TTS system. Instead, the aim is to produce a network that requires few data at deployment time to rapidly adapt to new speakers. We introduce and benchmark three strategies: (i) learning the speaker embedding while keeping the WaveNet core fixed, (ii) fine-tuning the entire architecture with stochastic gradient descent, and (iii) predicting the speaker embedding with a trained neural network encoder. The experiments show that these approaches are successful at adapting the multi-speaker neural network to new speakers, obtaining state-of-the-art results in both sample naturalness and voice similarity with merely a few minutes of audio data from new speakers.
Accepted by ICLR 2019
References in corpus (16)
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- WaveNet: A Generative Model for Raw Audio
- Pixel Recurrent Neural Networks
- Theoretical Models of Learning to Learn
- Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
- One-Shot Visual Imitation Learning via Meta-Learning
- Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron
- Neural Voice Cloning with a Few Samples
- Learning to Learn without Gradient Descent by Gradient Descent
- One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning
- Observe and Look Further: Achieving Consistent Performance on Atari
- Attentive Recurrent Comparators
- One-Shot Generalization in Deep Generative Models
- Online Learning with Gated Linear Networks
- Fitting New Speakers Based on a Short Untranscribed Sample
Cited by in corpus (25)
- Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
- Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder
- MelNet: A Generative Model for Audio in the Frequency Domain
- Weakly Supervised Disentanglement with Guarantees
- Meta-TTS: Meta-Learning for Few-Shot Speaker Adaptive Text-to-Speech
- Meta-learning of Sequential Strategies
- DAWSON: A Domain Adaptive Few Shot Generation Framework
- AdaDurIAN: Few-shot Adaptation for Neural Text-to-Speech with DurIAN
- Learning Compositional Neural Programs with Recursive Tree Search and Planning
- NAUTILUS: a Versatile Voice Cloning System
- Mel-spectrogram augmentation for sequence to sequence voice conversion
- Personal VAD: Speaker-Conditioned Voice Activity Detection
- A Unified Speaker Adaptation Method for Speech Synthesis using Transcribed and Untranscribed Speech with Backpropagation
- Modular Meta-Learning with Shrinkage
- Meta-Voice: Fast few-shot style transfer for expressive voice cloning using meta learning
- Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding
- Data Efficient Voice Cloning from Noisy Samples with Domain Adversarial Training
- CUHK-EE Voice Cloning System for ICASSP 2021 M2VoC Challenge
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- Using IPA-Based Tacotron for Data Efficient Cross-Lingual Speaker Adaptation and Pronunciation Enhancement
- AdaVocoder: Adaptive Vocoder for Custom Voice
- Continual Speaker Adaptation for Text-to-Speech Synthesis
- AdaSpeech 2: Adaptive Text to Speech with Untranscribed Data
- Human Languages in Source Code: Auto-Translation for Localized Instruction
- AI based Presentation Creator With Customized Audio Content Delivery