2 citations · 4 across the 3 of their papers we have counts for
4 papers · 1 filter
Towards zero-shot Text-based voice editing using acoustic context conditioning, utterance embeddings, and reference encoders
Jason Fong, Yun Wang, Prabhav Agrawal +4
Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edi…
Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling
Qing He, Zhiping Xiu, Thilo Koehler +1
Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates…
FBWave: Efficient and Scalable Neural Vocoders for Streaming Text-To-Speech on the Edge
Bichen Wu, Qing He, Peizhao Zhang +3
Nowadays more and more applications can benefit from edge-based text-to-speech (TTS). However, most existing TTS models are too computationally expensive and are not flexible enoug…
Interactive Text-to-Speech System via Joint Style Analysis
Yang Gao, Weiyi Zheng, Zhaojun Yang +3
While modern TTS technologies have made significant advancements in audio quality, there is still a lack of behavior naturalness compared to conversing with people. We propose a st…