14 papers
Andha-Dhun: A First Look at Audio Descriptions in Hindi
Ritabrata Chakraborty, Divy Kala, Nisheeth Bhooshan Gupta +3
Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum around ADs in movies and TV sho…
Attending to Multimodal Generation One Token at a Time
Varun Gupta, Vineet Gandhi, Makarand Tapaswi
Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving context. Prior work on interpretability h…
Steerable Visual Representations
Jona Ruthardt, Manu Gaur, Deva Ramanan +2
Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification,…
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
Darshan Singh, Zeeshan Khan, Makarand Tapaswi
Adapting CLIP for videos has gained popularity due to its semantic and rich representation. While CLIP is a good starting point, it typically undergoes post-pretraining (contrastiv…
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
Balaji Darur, Amanmeet Garg, Makarand Tapaswi
Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by re…
STRinGS: Selective Text Refinement in Gaussian Splatting
Abhinav Raundhal, Gaurav Behera, P J Narayanan +2
Text as signs, labels, or instructions is a critical element of real-world scenes as they can convey important contextual information. 3D representations such as 3D Gaussian Splatt…