Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
VIBE: Video Instruction-aligned Background music gEneration
Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj +5
Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representa…
cs.SD2026
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj +5
Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditi…