activity
20242026
collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2025

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation

Yakun Song, Jiawei Chen, Xiaobin Zhuang +9

Neural audio codecs have made significant strides in efficiently mapping raw audio waveforms into discrete token representations, which are foundational for contemporary audio gene…

cs.SD2025

Towards Reliable Large Audio Language Model

Ziyang Ma, Xiquan Li, Yakun Song +8

Recent advancements in large audio language models (LALMs) have demonstrated impressive results and promising prospects in universal understanding and reasoning across speech, musi…

cs.SD2025

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Ziyang Ma, Yinghao Ma, Yanqiao Zhu +31

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,00…

cs.SD2025

Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model

Ziyang Ma, Zhuo Chen, Yuping Wang +2

Large Audio-Language Models (LALMs) have demonstrated remarkable performance in tasks involving audio perception and understanding, such as speech recognition and audio captioning.…

cs.SD2025

Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Yi Yuan, Dongya Jia, Xiaobin Zhuang +9

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performan…