7 papers
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
Yan-Bo Lin, Jonah Casebeer, Long Mai +3
Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video…
A Generative-First Neural Audio Autoencoder
Jonah Casebeer, Ge Zhu, Zhepei Wang +1
Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single…
TAC: Timestamped Audio Captioning
Sonal Kumar, Prem Seetharaman, Ke Chen +8
Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introdu…
Stemphonic: All-at-once Flexible Multi-stem Music Generation
Shih-Lun Wu, Ge Zhu, Juan-Pablo Caceres +2
Music stem generation, the task of producing musically-synchronized and isolated instrument audio clips, offers the potential of greater user control and better alignment with musi…
Rethinking Music Captioning with Music Metadata LLMs
Irmak Bukey, Zhepei Wang, Chris Donahue +1
Music captioning, or the task of generating a natural language description of music, is useful for both music understanding and controllable music generation. Training captioning m…
DRAGON: Distributional Rewards Optimize Diffusion Generative Models
Yatong Bai, Jonah Casebeer, Somayeh Sojoudi +1
We present Distributional RewArds for Generative OptimizatioN (DRAGON), a versatile framework for fine-tuning media generation models towards a desired outcome. Compared with tradi…