2 papers
cs.SD2026
GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models
Zuyao You, Zhesong Yu, Mingyu Liu +3
In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamli…
cs.SD2024
MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
Hang Zhao, Yifei Xin, Zhesong Yu +3
In the realm of audio-language pre-training (ALP), the challenge of achieving cross-modal alignment is significant. Moreover, the integration of audio inputs with diverse distribut…