24 citations · 55 across the 5 of their papers we have counts for
5 papers · 1 filter
Seeing the Pose in the Pixels: Learning Pose-Aware Representations in Vision Transformers
Dominick Reilly, Aman Chadha, Srijan Das
Human perception of surroundings is often guided by the various poses present within the environment. Many computer vision tasks, such as human action recognition and robot imitati…
I see what you hear: a vision-inspired method to localize words
Mohammad Samragh, Arnav Kundu, Ting-Yao Hu +5
This paper explores the possibility of using visual object detection techniques for word localization in speech data. Object detection has been thoroughly studied in the contempora…
iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability
Aman Chadha, Vinija Jain
Causality knowledge is vital to building robust AI systems. Deep learning models often perform poorly on tasks that require causal reasoning, which is often derived using some form…
iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering
Aman Chadha, Gurneet Arora, Navpreet Kaloty
Most prior art in visual understanding relies solely on analyzing the "what" (e.g., event recognition) and "where" (e.g., event localization), which in some cases, fails to describ…
iSeeBetter: Spatio-temporal video super-resolution using recurrent generative back-projection networks
Aman Chadha, John Britto, M. Mani Roja
Recently, learning-based models have enhanced the performance of single-image super-resolution (SISR). However, applying SISR successively to each video frame leads to a lack of te…