activity
20172026
most citedSituation Recognition with Graph Neural Networks

22 citations · 67 across the 20 of their papers we have counts for

collaborators
Showing cs.CVShow all

33 papers · 1 filter

cs.CV2026

Andha-Dhun: A First Look at Audio Descriptions in Hindi

Ritabrata Chakraborty, Divy Kala, Nisheeth Bhooshan Gupta +3

Audio Descriptions (ADs) narrate visual content for Blind and Low Vision (BLV) audiences during gaps in audiovisual media. There is growing momentum around ADs in movies and TV sho…

cs.CV2026

Attending to Multimodal Generation One Token at a Time

Varun Gupta, Vineet Gandhi, Makarand Tapaswi

Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving context. Prior work on interpretability h…

cs.CV2026

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

Balaji Darur, Amanmeet Garg, Makarand Tapaswi

Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by re…

cs.CV2026

Steerable Visual Representations

Jona Ruthardt, Manu Gaur, Deva Ramanan +2

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification,…

cs.CV2025

STRinGS: Selective Text Refinement in Gaussian Splatting

Abhinav Raundhal, Gaurav Behera, P J Narayanan +2

Text as signs, labels, or instructions is a critical element of real-world scenes as they can convey important contextual information. 3D representations such as 3D Gaussian Splatt…

cs.CV2025

MALeR: Improving Compositional Fidelity in Layout-Guided Generation

Shivank Saxena, Dhruv Srivastava, Makarand Tapaswi

Recent advances in text-to-image models have enabled a new era of creative and controllable image generation. However, generating compositional scenes with multiple subjects and at…