8 papers
Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
Haven Kim, Zachary Novack, Julian McAuley +1
Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reprodu…
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
Zachary Novack, Stephen Brade, Haven Kim +8
Interactive streaming music generation promises the use of generative models for live performance and co-creation that is impossible with offline models. However, SOTA models exist…
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
Haven Kim, Julian McAuley
Conversational music recommendation (CMR) research currently faces a tradeoff between authentic dialogue corpora that are limited in scale and synthesized corpora that scale up but…
Expressiveness Limits of Autoregressive Semantic ID Generation in Generative Recommendation
Yupeng Hou, Haven Kim, Clark Mingxuan Ju +3
Generative recommendation (GR) models generate items by autoregressively producing a sequence of discrete tokens that jointly index the target item. However, this autoregressive ge…
A Design Space for Live Music Agents
Yewon Kim, Stephen Brade, Alexander Wang +7
Live music provides a uniquely rich setting for studying creativity and interaction due to its spontaneous nature. The pursuit of live music agents--intelligent systems supporting…
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
Haven Kim, Yupeng Hou, Julian McAuley
Generative recommendation systems have achieved significant advances by leveraging semantic IDs to represent items. However, existing approaches that tokenize each modality indepen…