collaborators

14 papers

cs.CV2026

Andha-Dhun: A First Look at Audio Descriptions in Hindi

Ritabrata Chakraborty, Divy Kala, Nisheeth Bhooshan Gupta +3

Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum around ADs in movies and TV sho…

cs.CV2026

Attending to Multimodal Generation One Token at a Time

Varun Gupta, Vineet Gandhi, Makarand Tapaswi

Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving context. Prior work on interpretability h…

cs.CV2026

Steerable Visual Representations

Jona Ruthardt, Manu Gaur, Deva Ramanan +2

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification,…

cs.CV2026

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels

Darshan Singh, Zeeshan Khan, Makarand Tapaswi

Adapting CLIP for videos has gained popularity due to its semantic and rich representation. While CLIP is a good starting point, it typically undergoes post-pretraining (contrastiv…

cs.CV2026

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

Balaji Darur, Amanmeet Garg, Makarand Tapaswi

Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by re…

cs.CV2025

STRinGS: Selective Text Refinement in Gaussian Splatting

Abhinav Raundhal, Gaurav Behera, P J Narayanan +2

Text as signs, labels, or instructions is a critical element of real-world scenes as they can convey important contextual information. 3D representations such as 3D Gaussian Splatt…