1 paper · 1 filter
Wenhao Zou, Yuwei Miao, Zhanyu Ma +5
Real-time voice agents face a dilemma: end-to-end models often lack deep reasoning, while cascaded pipelines incur high latency by executing ASR, LLM reasoning, and TTS strictly in…