119 citations · 210 across the 25 of their papers we have counts for
38 papers · 1 filter
ClipSitu: Effectively Leveraging CLIP for Conditional Predictions in Situation Recognition
Debaditya Roy, Dhruv Verma, Basura Fernando
Situation Recognition is the task of generating a structured summary of what is happening in an image using an activity verb and the semantic roles played by actors and objects. In…
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
Ishaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal +2
While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the…
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
Ramanathan Rajendiran, Debaditya Roy, Basura Fernando
Humans have the natural ability to recognize actions even if the objects involved in the action or the background are changed. Humans can abstract away the action from the appearan…
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
Hao Zhang, Yeo Keat Ee, Basura Fernando
Visual abductive reasoning aims to make likely explanations for visual observations. We propose a simple yet effective Region Conditioned Adaptation, a hybrid parameter-efficient f…
Who are you referring to? Coreference resolution in image narrations
Arushi Goel, Basura Fernando, Frank Keller +1
Coreference resolution aims to identify words and phrases which refer to same entity in a text, a core task in natural language processing. In this paper, we extend this task to re…
Interaction Region Visual Transformer for Egocentric Action Anticipation
Debaditya Roy, Ramanathan Rajendiran, Basura Fernando
Human-object interaction is one of the most important visual cues and we propose a novel way to represent human-object interactions for egocentric action anticipation. We propose a…