2 citations · 3 across the 2 of their papers we have counts for
2 papers
eess.AS2023★ 1 cited
Speak While You Think: Streaming Speech Synthesis During Text Generation
Avihu Dekel, Slava Shechtman, Raul Fernandez +3
Large Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outpu…
eess.AS2022★ 2 cited
Extending RNN-T-based speech recognition systems with emotion and language classification
Zvi Kons, Hagai Aronowitz, Edmilson Morais +4
Speech transcription, emotion recognition, and language identification are usually considered to be three different tasks. Each one requires a different model with a different arch…