184 citations · 367 across the 11 of their papers we have counts for
4 papers · 1 filter
Parallel Tacotron: Non-Autoregressive and Controllable TTS
Isaac Elias, Heiga Zen, Jonathan Shen +4
Although neural end-to-end text-to-speech models can synthesize highly natural speech, there is still room for improvements to its efficiency and naturalness. This paper proposes a…
Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling
Jonathan Shen, Ye Jia, Mike Chrzanowski +4
This paper presents Non-Attentive Tacotron based on the Tacotron 2 text-to-speech model, replacing the attention mechanism with an explicit duration predictor. This improves robust…
Textual Echo Cancellation
Shaojin Ding, Ye Jia, Ke Hu +1
In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can…
Improved Noisy Student Training for Automatic Speech Recognition
Daniel S. Park, Yu Zhang, Ye Jia +5
Recently, a semi-supervised learning method known as "noisy student training" has been shown to improve image classification performance of deep networks significantly. Noisy stude…