1 paper · 1 filter
Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj +5
Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representa…