2 papers
cs.CV2025
Investigating Mechanisms for In-Context Vision Language Binding
Darshana Saravanan, Makarand Tapaswi, Vineet Gandhi
To understand a prompt, Vision-Language models (VLMs) must perceive the image, comprehend the text, and build associations within and across both modalities. For instance, given an…
cs.CV2024
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
Darshana Saravanan, Varun Gupta, Darshan Singh +3
A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision or v…