activity
20202023
most citediPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering

24 citations · 55 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20234 cited

Seeing the Pose in the Pixels: Learning Pose-Aware Representations in Vision Transformers

Dominick Reilly, Aman Chadha, Srijan Das

Human perception of surroundings is often guided by the various poses present within the environment. Many computer vision tasks, such as human action recognition and robot imitati…

cs.CV2022

I see what you hear: a vision-inspired method to localize words

Mohammad Samragh, Arnav Kundu, Ting-Yao Hu +5

This paper explores the possibility of using visual object detection techniques for word localization in speech data. Object detection has been thoroughly studied in the contempora…

cs.CV20214 cited

iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability

Aman Chadha, Vinija Jain

Causality knowledge is vital to building robust AI systems. Deep learning models often perform poorly on tasks that require causal reasoning, which is often derived using some form…

cs.CV202024 cited

iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering

Aman Chadha, Gurneet Arora, Navpreet Kaloty

Most prior art in visual understanding relies solely on analyzing the "what" (e.g., event recognition) and "where" (e.g., event localization), which in some cases, fails to describ…

cs.CV202023 cited

iSeeBetter: Spatio-temporal video super-resolution using recurrent generative back-projection networks

Aman Chadha, John Britto, M. Mani Roja

Recently, learning-based models have enhanced the performance of single-image super-resolution (SISR). However, applying SISR successively to each video frame leads to a lack of te…