2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CL2024
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
Maohao Shen, Shun Zhang, Jilong Wu +5
Large language models (LLMs) have revolutionized natural language processing (NLP) with impressive performance across various text-based tasks. However, the extension of text-domin…
eess.AS2024
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
Jinzheng Zhao, Niko Moritz, Egor Lakomkin +7
Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumu…
cs.SD2021★ 2 cited
Multi-rate attention architecture for fast streamable Text-to-speech spectrum modeling
Qing He, Zhiping Xiu, Thilo Koehler +1
Typical high quality text-to-speech (TTS) systems today use a two-stage architecture, with a spectrum model stage that generates spectral frames and a vocoder stage that generates…