19 citations · 24 across the 3 of their papers we have counts for
5 papers
YORO -- Lightweight End to End Visual Grounding
Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani +2
We present YORO - a multi-modal transformer encoder-only architecture for the Visual Grounding (VG) task. This task involves localizing, in an image, an object referred via natural…
DocFormer: End-to-End Transformer for Document Understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota +2
We present DocFormer -- a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand docu…
Skeleton based Zero Shot Action Recognition in Joint Pose-Language Semantic Space
Bhavan Jasani, Afshaan Mazagonwalla
How does one represent an action? How does one describe an action that we have never seen before? Such questions are addressed by the Zero Shot Learning paradigm, where a model is…
Are we asking the right questions in MovieQA?
Bhavan Jasani, Rohit Girdhar, Deva Ramanan
Joint vision and language tasks like visual question answering are fascinating because they explore high-level understanding, but at the same time, can be more prone to language bi…
Learning Sampling Policies for Domain Adaptation
Yash Patel, Kashyap Chitta, Bhavan Jasani
We address the problem of semi-supervised domain adaptation of classification algorithms through deep Q-learning. The core idea is to consider the predictions of a source domain ne…