activity
20162026
most citedRecovering the Missing Link: Predicting Class-Attribute Associations for Unsupervised Zero-Shot Learning

22 citations · 105 across the 37 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

8 papers · 2 filters

cs.CV2024

The Sound of Water: Inferring Physical Properties from Pouring Liquids

Piyush Bagad, Makarand Tapaswi, Cees G. M. Snoek +1

We study the connection between audio-visual observations and the underlying physics of a mundane yet intriguing everyday activity: pouring liquids. Given only the sound of liquid…

cs.CV2024

Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation

Manu Gaur, Darshan Singh S, Makarand Tapaswi

Visual Question Answering (VQA) with multiple choice questions enables a vision-centric evaluation of Multimodal Large Language Models (MLLMs). Although it reliably checks the exis…

cs.CV2024

No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning

Manu Gaur, Darshan Singh, Makarand Tapaswi

Image captioning systems are unable to generate fine-grained captions as they are trained on data that is either noisy (alt-text) or generic (human annotations). This is further ex…

cs.CV2024★ 1 cited

VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment

Darshana Saravanan, Varun Gupta, Darshan Singh +3

A fundamental aspect of compositional reasoning in a video is associating people and their actions across time. Recent years have seen great progress in general-purpose vision or v…

cs.CV2024

"Previously on ..." From Recaps to Story Summarization

Aditya Kumar Singh, Dhruv Srivastava, Makarand Tapaswi

We introduce multimodal story summarization by leveraging TV episode recaps - short video sequences interweaving key story moments from previous episodes to bring viewers up to spe…

cs.CV2024

MICap: A Unified Model for Identity-aware Movie Descriptions

Haran Raajesh, Naveen Reddy Desanur, Zeeshan Khan +1

Characters are an important aspect of any storyline and identifying and including them in descriptions is necessary for story understanding. While previous work has largely ignored…