1 paper · 1 filter
Jiaming Zhou, Xuxin Cheng, Shiwan Zhao +5
Autoregressive (AR) large audio language models (LALMs) such as Qwen-2.5-Omni have achieved strong performance on audio understanding and interaction, but scaling them remains cost…