Sprachsynthese -- State-of-the-Art in englischer und deutscher Sprache
arXiv:2106.06230
Abstract
Reading text aloud is an important feature for modern computer applications. It not only facilitates access to information for visually impaired people, but is also a pleasant convenience for non-impaired users. In this article, the state of the art of speech synthesis is presented separately for mel-spectrogram generation and vocoders. It concludes with an overview of available data sets for English and German with a discussion of the transferability of the good speech synthesis results from English to German language.
in German
References in corpus (5)
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
- Parallel Tacotron: Non-Autoregressive and Controllable TTS
- TalkNet: Fully-Convolutional Non-Autoregressive Speech Synthesis Model