9 citations · 19 across the 9 of their papers we have counts for
Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR
Zhiwei Lin, Kaiqi Fu, Rime Wen +5
Streaming automatic speech recognition (ASR) for real-time voice agents and full-duplex dialogue must provide accurate partial transcripts with low commit latency. Existing systems…
cs.SD2025★ 1 cited
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Zhihao Du, Changfeng Gao, Yuxuan Wang +19
In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model…
cs.SD2024★ 9 cited
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Keyu An, Qian Chen, Chong Deng +30
This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative mo…