1 paper
Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj +5
Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representa…