2 papers
cs.CV2026
Attending to Multimodal Generation One Token at a Time
Varun Gupta, Vineet Gandhi, Makarand Tapaswi
Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving context. Prior work on interpretability h…
cs.CV2025
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
Darshana Saravanan, Varun Gupta, Darshan Singh +3
A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision or v…