8 papers · 1 filter
VIBE: Video Instruction-aligned Background music gEneration
Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj +5
Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representa…
TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3
Large audio-language models (LALMs) describe audio at the clip level but cannot assign timestamps to the events, speakers, or sounds they identify. Despite being essential for down…
DuplexWorld: Can voice agents help you get through the day?
Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli +3
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversatio…
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji +1
Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered…
FIGMA: Towards FIne-Grained Music retrievAl
Nishit Anand, Ashish Seth, Sreyan Ghosh +2
Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries. Wh…
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj +5
Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditi…