8 papers
Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio
Avi Gupta, Nilotpal Sinha, Vishnu Raj +4
Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have spurred significant interes…
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj +5
Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditi…
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
Xiyang Wu, Zongxia Li, Jihui Jin +7
Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial interactions. We present a nove…
CompanionCast: Toward Social Collaboration with Multi-Agent Systems in Shared Experiences
Yiyang Wang, Chen Chen, Tica Lin +4
Shared experiences are fundamental to social connection, yet media consumption is increasingly solitary. While AI companions offer real-time reactions and emotional regulation, exi…
Glow with the Flow: AI-Assisted Creation of Ambient Lightscapes for Music Videos
Frederic Anthony Robinson, Vishnu Raj, David Cooper +2
Designed light is an established modality for live performance and music playback. Despite the growing availability of consumer smart lighting, the creation of designed light for m…
Enhanced Generative Machine Listener
Vishnu Raj, Gouthaman KV, Shiv Gehlot +2
We present GMLv2, a reference-based model designed for the prediction of subjective audio quality as measured by MUSHRA scores. GMLv2 introduces a Beta distribution-based loss to m…