11 papers
Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
Haven Kim, Zachary Novack, Julian McAuley +1
Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reprodu…
Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning
Hansen Jin Lillemark, Alex Rojas, Zachary Novack +5
Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting inference procedu…
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
Phillip Long, Zachary Novack, Chris Donahue
Autoregressive "language" models (LMs) trained on raw waveforms can be repurposed for lossless audio compression, but prior work is limited to 8-bit audio, leaving open whether suc…
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
Zachary Novack, Stephen Brade, Haven Kim +8
Interactive streaming music generation promises the use of generative models for live performance and co-creation that is impossible with offline models. However, SOTA models exist…
Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis
Shuyang Cui, Zhi Zhong, Qiyu Wu +9
Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial efforts of creators. While recent g…
Low-Resource Guidance for Controllable Latent Audio Diffusion
Zachary Novack, Zack Zukowski, CJ Carr +6
Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guid…