activity
20242026
collaborators

9 papers

cs.SD2026

MusPyExpress: Extending MusPy with Enhanced Expression Text Support

Phillip Long, Hao-Wen Dong, Julian McAuley +1

Current work in modeling symbolic music primarily relies on representations extracted from MIDI-like data. While such formats allow for modeling symbolic music as sequences of note…

cs.SD2026

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

Haven Kim, Zachary Novack, Julian McAuley +1

Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reprodu…

cs.SD2026

Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods

Fang-Chih Hsieh, Wei-Jaw Lee, Chun-Ping Wang +3

This paper presents an overview and the technical framework of the ICME 2026 Grand Challenge on Academic Text-to-Music Generation (ATTM). Despite the rapid progress in text-to-musi…

eess.AS2026

Unmixing The Crowd: Learning Persistent Speaker Representations from Mixture-Derived Multi-Speaker Embeddings

Sidharth Sidharth, Meysam Asgari, Hao-Wen Dong +1

We study whether persistent conversational speaker structure can be extracted directly from local overlapping speech mixtures. We propose a teacher-student framework that learns mi…

cs.SD2025

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Haven Kim, Zachary Novack, Weihan Xu +2

Despite recent advancements in music generation systems, their application in film production remains limited, as they struggle to capture the nuances of real-world filmmaking, whe…

cs.AI2025

Synthesizing Composite Hierarchical Structure from Symbolic Music Corpora

Ilana Shapiro, Ruanqianqian Huang, Zachary Novack +5

Western music is an innately hierarchical system of interacting levels of structure, from fine-grained melody to high-level form. In order to analyze music compositions holisticall…