9 papers
MusPyExpress: Extending MusPy with Enhanced Expression Text Support
Phillip Long, Hao-Wen Dong, Julian McAuley +1
Current work in modeling symbolic music primarily relies on representations extracted from MIDI-like data. While such formats allow for modeling symbolic music as sequences of note…
Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
Haven Kim, Zachary Novack, Julian McAuley +1
Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reprodu…
Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods
Fang-Chih Hsieh, Wei-Jaw Lee, Chun-Ping Wang +3
This paper presents an overview and the technical framework of the ICME 2026 Grand Challenge on Academic Text-to-Music Generation (ATTM). Despite the rapid progress in text-to-musi…
Unmixing The Crowd: Learning Persistent Speaker Representations from Mixture-Derived Multi-Speaker Embeddings
Sidharth Sidharth, Meysam Asgari, Hao-Wen Dong +1
We study whether persistent conversational speaker structure can be extracted directly from local overlapping speech mixtures. We propose a teacher-student framework that learns mi…
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
Haven Kim, Zachary Novack, Weihan Xu +2
Despite recent advancements in music generation systems, their application in film production remains limited, as they struggle to capture the nuances of real-world filmmaking, whe…
Synthesizing Composite Hierarchical Structure from Symbolic Music Corpora
Ilana Shapiro, Ruanqianqian Huang, Zachary Novack +5
Western music is an innately hierarchical system of interacting levels of structure, from fine-grained melody to high-level form. In order to analyze music compositions holisticall…