3 papers
eess.AS2026
AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation
Huadai Liu, Kaicheng Luo, Wen Wang +4
Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis tha…
cs.SD2025
InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
Chong Zhang, Yukun Ma, Qian Chen +12
We introduce InspireMusic, a framework integrated super resolution and large language model for high-fidelity long-form music generation. A unified framework generates high-fidelit…
cs.CL2025
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Qian Chen, Yafeng Chen, Yanni Chen +33
Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and hum…