15 papers
Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
Haven Kim, Zachary Novack, Julian McAuley +1
Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reprodu…
Retrieval Augmented Conversational Recommendation with Reinforcement Learning
Zhenrui Yue, Honglei Zhuang, Zhen Qin +4
Large language models (LLMs) exhibit enhanced capabilities in language understanding and generation. By utilizing their embedded knowledge, LLMs are increasingly used as conversati…
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
Lei Zhang, Junjiao Tian, Zhipeng Fan +9
Humans paint images incrementally: they plan a global layout, sketch a coarse draft, inspect, and refine details, and most importantly, each step is grounded in the evolving visual…
Steering Autoregressive Music Generation with Recursive Feature Machines
Daniel Zhao, Daniel Beaglehole, Taylor Berg-Kirkpatrick +2
Controllable music generation remains a significant challenge, with existing methods often requiring model retraining or introducing audible artifacts. We introduce MusicRFM, a fra…
MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
Jingyue Huang, Zachary Novack, Phillip Long +4
Discrete representation learning has shown promising results across various domains, including generation and understanding in image, speech and language. Inspired by these advance…
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
Qihui Yang, Taylor Berg-Kirkpatrick, Julian McAuley +1
Despite rapid progress in end-to-end AI music generation, AI-driven modeling of professional Digital Signal Processing (DSP) workflows remains challenging. In particular, while the…