activity
20182024
most citedSkeleton based Zero Shot Action Recognition in Joint Pose-Language Semantic Space

19 citations · 24 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2024

Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA

Zhuowan Li, Bhavan Jasani, Peng Tang +1

Understanding data visualizations like charts and plots requires reasoning about both visual elements and numerics. Although strong in extractive questions, current chart visual qu…

cs.CV2022★ 1 cited

YORO -- Lightweight End to End Visual Grounding

Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani +2

We present YORO - a multi-modal transformer encoder-only architecture for the Visual Grounding (VG) task. This task involves localizing, in an image, an object referred via natural…

cs.CV2021★ 4 cited

DocFormer: End-to-End Transformer for Document Understanding

Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota +2

We present DocFormer -- a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand docu…

cs.CV2019★ 19 cited

Skeleton based Zero Shot Action Recognition in Joint Pose-Language Semantic Space

Bhavan Jasani, Afshaan Mazagonwalla

How does one represent an action? How does one describe an action that we have never seen before? Such questions are addressed by the Zero Shot Learning paradigm, where a model is…

cs.CV2019

Are we asking the right questions in MovieQA?

Bhavan Jasani, Rohit Girdhar, Deva Ramanan

Joint vision and language tasks like visual question answering are fascinating because they explore high-level understanding, but at the same time, can be more prone to language bi…

cs.CV2018

Learning Sampling Policies for Domain Adaptation

Yash Patel, Kashyap Chitta, Bhavan Jasani

We address the problem of semi-supervised domain adaptation of classification algorithms through deep Q-learning. The core idea is to consider the predictions of a source domain ne…