6 papers
Native Audio-Visual Alignment for Generation
Longbin Ji, Guan Wang, Xuan Wei +6
Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual-acoustic content. However, existing open-source methods mainly rely on eith…
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
Ziqi Zhao, Xinyu Ma, Liu Yang +6
On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-policy rollouts. However, ex…
Eureka-Audio: Triggering Audio Intelligence in Compact Language Models
Dan Zhang, Yishu Lei, Jing Hu +10
We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range…
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
Jing Hu, Danxiang Zhu, Xianlong Luo +9
Large Audio Language Models (LALMs) have garnered significant research interest. Despite being built upon text-based large language models (LLMs), LALMs frequently exhibit a degrad…
MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free
Yishu Lei, Shuwei He, Jing Hu +9
Extending the input modality of Large Language Models~(LLMs) to the audio domain is essential for achieving comprehensive multimodal perception. However, it is well-known that acou…
MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
Hongwei Chen, Yishu Lei, Dan Zhang +10
Test-time scaling has emerged as a promising paradigm in language modeling, wherein additional computational resources are allocated during inference to enhance model performance.…