2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.SD2021★ 2 cited
Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling
Qing He, Zhiping Xiu, Thilo Koehler +1
Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates…
eess.AS2019★ 1 cited
G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR
Duc Le, Thilo Koehler, Christian Fuegen +1
Grapheme-based acoustic modeling has recently been shown to outperform phoneme-based approaches in both hybrid and end-to-end automatic speech recognition (ASR), even on non-phonem…