20 citations · 29 across the 9 of their papers we have counts for
13 papers · 1 filter
Text Spotting Transformers
Xiang Zhang, Yongwen Su, Subarna Tripathi +1
In this paper, we present TExt Spotting TRansformers (TESTR), a generic end-to-end text spotting framework using Transformers for text detection and recognition in the wild. TESTR…
Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar +1
We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and…
Learning of Visual Relations: The Devil is in the Tails
Alakh Desai, Tz-Ying Wu, Subarna Tripathi +1
Significant effort has been recently devoted to modeling visual relations. This has mostly addressed the design of architectures, typically by adding parameters and increasing mode…
In Defense of Scene Graphs for Image Captioning
Kien Nguyen, Subarna Tripathi, Bang Du +2
The mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models. Recently, image scene graphs have been u…
Structured Query-Based Image Retrieval Using Scene Graphs
Brigit Schroeder, Subarna Tripathi
A structured query can capture the complexity of object interactions (e.g. 'woman rides motorcycle') unlike single objects (e.g. 'woman' or 'motorcycle'). Retrieval using structure…
Triplet-Aware Scene Graph Embeddings
Brigit Schroeder, Subarna Tripathi, Hanlin Tang
Scene graphs have become an important form of structured knowledge for tasks such as for image generation, visual relation detection, visual question answering, and image retrieval…