GANtron: Emotional Speech Synthesis with Generative Adversarial Networks
arXiv:2110.03390
Abstract
Speech synthesis is used in a wide variety of industries. Nonetheless, it always sounds flat or robotic. The state of the art methods that allow for prosody control are very cumbersome to use and do not allow easy tuning. To tackle some of these drawbacks, in this work we target the implementation of a text-to-speech model where the inferred speech can be tuned with the desired emotions. To do so, we use Generative Adversarial Networks (GANs) together with a sequence-to-sequence model using an attention mechanism. We evaluate four different configurations considering different inputs and training strategies, study them and prove how our best model can generate speech files that lie in the same distribution as the initial training dataset. Additionally, a new strategy to boost the training convergence by applying a guided attention loss is proposed.
9 pages, 4 figures
References in corpus (6)
- WaveNet: A Generative Model for Raw Audio
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Deep Voice: Real-time Neural Text-to-Speech
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
- Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
- Uncovering Latent Style Factors for Expressive Speech Synthesis