3 citations · 3 across the 2 of their papers we have counts for
2 papers
eess.AS2025
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
Wenrui Liu, Qian Chen, Wen Wang +11
Neural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neu…
cs.CL2025★ 3 cited
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Qian Chen, Yafeng Chen, Yanni Chen +33
Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and hum…