4 papers
Towards Effective Negation Modeling in Joint Audio-Text Models for Music
Yannis Vasilakis, Rachel Bittner, Johan Pauwels
Joint audio-text models are widely used for music retrieval, yet they struggle with semantic phenomena such as negation. Negation is fundamental for distinguishing the absence (or…
Evaluation of pretrained language models on music understanding
Yannis Vasilakis, Rachel Bittner, Johan Pauwels
Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song genera…
I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition
Yannis Vasilakis, Rachel Bittner, Johan Pauwels
Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels. These…
LLark: A Multimodal Instruction-Following Language Model for Music
Josh Gardner, Simon Durand, Daniel Stoller +1
Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms o…