collaborators

8 papers

cs.CV2026

Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio

Avi Gupta, Nilotpal Sinha, Vishnu Raj +4

Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have spurred significant interes…

cs.SD2026

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj +5

Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditi…

cs.CV2026

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

Xiyang Wu, Zongxia Li, Jihui Jin +7

Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial interactions. We present a nove…

cs.HC2026

CompanionCast: Toward Social Collaboration with Multi-Agent Systems in Shared Experiences

Yiyang Wang, Chen Chen, Tica Lin +4

Shared experiences are fundamental to social connection, yet media consumption is increasingly solitary. While AI companions offer real-time reactions and emotional regulation, exi…

cs.HC2026

Glow with the Flow: AI-Assisted Creation of Ambient Lightscapes for Music Videos

Frederic Anthony Robinson, Vishnu Raj, David Cooper +2

Designed light is an established modality for live performance and music playback. Despite the growing availability of consumer smart lighting, the creation of designed light for m…

eess.AS2026

Enhanced Generative Machine Listener

Vishnu Raj, Gouthaman KV, Shiv Gehlot +2

We present GMLv2, a reference-based model designed for the prediction of subjective audio quality as measured by MUSHRA scores. GMLv2 introduces a Beta distribution-based loss to m…