activity
20242026
collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2026

Assessing Factual Music Comprehension in Large Audio Language Models

Daniel Chenyu Lin, Michael Freeman, John Thickstun

Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empiri…

cs.SD2026

Music Transcription with (Almost) No Supervision

Saebyeol Shin, Chao Wan, Zhenzhen Liu +4

Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions.…

cs.SD2025

Robust Neural Audio Fingerprinting using Music Foundation Models

Shubhr Singh, Kiran Bhat, Xavier Riley +3

The proliferation of distorted, compressed, and manipulated music on modern media platforms like TikTok motivates the development of more robust audio fingerprinting techniques to…

cs.SD2025

Aligning Text-to-Music Evaluation with Human Preferences

Yichen Huang, Zachary Novack, Koichi Saito +5

Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fréche…

cs.SD2025

Hookpad Aria: A Copilot for Songwriters

Chris Donahue, Shih-Lun Wu, Yewon Kim +3

We present Hookpad Aria, a generative AI system designed to assist musicians in writing Western pop songs. Our system is seamlessly integrated into Hookpad, a web-based editor desi…

cs.SD2024

Anticipatory Music Transformer

John Thickstun, David Hall, Chris Donahue +1

We introduce anticipation: a method for constructing a controllable generative model of a temporal point process (the event process) conditioned asynchronously on realizations of a…