4 papers
Unifying Score and Performance for Fine-Grained Music Understanding in Audio-Language Models
Milan Liessens Dujardin, Song-Ze Yu, Kevin Miao
Large audio language models (LALMs) have shown promising progress in broad music-understanding tasks such as tagging, retrieval, and captioning. Music understanding that requires f…
MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance
Milan Liessens Dujardin, Song-Ze Yu, Kevin Miao
Musicians commonly communicate music through scores and performances. Scores encode musical intent, while performances realize it in sound. To investigate whether models can meanin…
InstructFX2FX: A Multi-Turn Text-to-Effect System for Sequential Audio Effect Refinement
Song-Ze Yu, Milan Liessens Dujardin, Yuxuan Cai +4
We present InstructFX2FX, a system for sequential audio effect refinement through multi-turn natural-language instructions. Existing text-to-effect systems are largely single-shot,…
PitchBench: Measuring Pitch Hearing in Audio-Language Models
Milan Liessens Dujardin, Song-Ze Yu, Craver Corbyn Thomas-Smith +2
Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription to captioning, recommendation…