4 citations · 8 across the 7 of their papers we have counts for
7 papers
GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models
Zuyao You, Zhesong Yu, Mingyu Liu +3
In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamli…
MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
Hang Zhao, Yifei Xin, Zhesong Yu +3
In the realm of audio-language pre-training (ALP), the challenge of achieving cross-modal alignment is significant. Moreover, the integration of audio inputs with diverse distribut…
Joint Music and Language Attention Models for Zero-shot Music Tagging
Xingjian Du, Zhesong Yu, Jiaju Lin +2
Music tagging is a task to predict the tags of music recordings. However, previous music tagging research primarily focuses on close-set music tagging tasks which can not be genera…
Attention-based cross-modal fusion for audio-visual voice activity detection in musical video streams
Yuanbo Hou, Zhesong Yu, Xia Liang +4
Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse mu…
Contrastive Unsupervised Learning for Audio Fingerprinting
Zhesong Yu, Xingjian Du, Bilei Zhu +1
The rise of video-sharing platforms has attracted more and more people to shoot videos and upload them to the Internet. These videos mostly contain a carefully-edited background au…
ByteCover: Cover Song Identification via Multi-Loss Training
Xingjian Du, Zhesong Yu, Bilei Zhu +2
We present in this paper ByteCover, which is a new feature learning method for cover song identification (CSI). ByteCover is built based on the classical ResNet model, and two majo…