collaborators

13 papers

cs.SD2026

DuplexWorld: Can voice agents help you get through the day?

Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli +3

Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversatio…

cs.SD2026

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji +1

Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered…

eess.AS2026

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar +19

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, a…

cs.SD2026

FIGMA: Towards FIne-Grained Music retrievAl

Nishit Anand, Ashish Seth, Sreyan Ghosh +2

Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries. Wh…

cs.AI2026

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models

Rayyan Abdalla, Amir Hussein, Min Wu +1

Post-training quantization (PTQ) is critical for the efficient deployment of large language models (LLMs). Recent ultra-low-bit PTQ methods rely on rigid weight-saliency assumption…

cs.SD2026

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj +5

Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditi…