91 citations · 97 across the 5 of their papers we have counts for
10 papers · 1 filter
DocTr: Document Transformer for Structured Information Extraction in Documents
Haofu Liao, Aruni RoyChowdhury, Weijian Li +6
We present a new formulation for structured information extraction (SIE) from visually rich documents. It aims to address the limitations of existing IOB tagging or graph-based for…
Pose And Joint-Aware Action Recognition
Anshul Shah, Shlok Mishra, Ankan Bansal +3
Recent progress on action recognition has mainly focused on RGB and optical flow features. In this paper, we approach the problem of joint-based action recognition. Unlike other mo…
Visual Question Answering on Image Sets
Ankan Bansal, Yuting Zhang, Rama Chellappa
We introduce the task of Image-Set Visual Question Answering (ISVQA), which generalizes the commonly studied single-image VQA problem to multi-image settings. Taking a natural lang…
Spatial Priming for Detecting Human-Object Interactions
Ankan Bansal, Sai Saketh Rambhatla, Abhinav Shrivastava +1
The relative spatial layout of a human and an object is an important cue for determining how they interact. However, until now, spatial layout has been used just as side-informatio…
How are attributes expressed in face DCNNs?
Prithviraj Dhar, Ankan Bansal, Carlos D. Castillo +3
As deep networks become increasingly accurate at recognizing faces, it is vital to understand how these networks process faces. While these networks are solely trained to recognize…
Detecting Human-Object Interactions via Functional Generalization
Ankan Bansal, Sai Saketh Rambhatla, Abhinav Shrivastava +1
We present an approach for detecting human-object interactions (HOIs) in images, based on the idea that humans interact with functionally similar objects in a similar manner. The p…