19 papers
MusPyExpress: Extending MusPy with Enhanced Expression Text Support
Phillip Long, Hao-Wen Dong, Julian McAuley +1
Current work in modeling symbolic music primarily relies on representations extracted from MIDI-like data. While such formats allow for modeling symbolic music as sequences of note…
Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
Haven Kim, Zachary Novack, Julian McAuley +1
Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reprodu…
Retrieval Augmented Conversational Recommendation with Reinforcement Learning
Zhenrui Yue, Honglei Zhuang, Zhen Qin +4
Large language models (LLMs) exhibit enhanced capabilities in language understanding and generation. By utilizing their embedded knowledge, LLMs are increasingly used as conversati…
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
Lei Zhang, Junjiao Tian, Zhipeng Fan +9
Humans paint images incrementally: they plan a global layout, sketch a coarse draft, inspect, and refine details, and most importantly, each step is grounded in the evolving visual…
Steering Autoregressive Music Generation with Recursive Feature Machines
Daniel Zhao, Daniel Beaglehole, Taylor Berg-Kirkpatrick +2
Controllable music generation remains a significant challenge, with existing methods often requiring model retraining or introducing audible artifacts. We introduce MusicRFM, a fra…
MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
Jingyue Huang, Zachary Novack, Phillip Long +4
Discrete representation learning has shown promising results across various domains, including generation and understanding in image, speech and language. Inspired by these advance…