115 citations · 157 across the 6 of their papers we have counts for
10 papers · 1 filter
YORO -- Lightweight End to End Visual Grounding
Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani +2
We present YORO - a multi-modal transformer encoder-only architecture for the Visual Grounding (VG) task. This task involves localizing, in an image, an object referred via natural…
Towards Weakly-Supervised Text Spotting using a Multi-Task Transformer
Yair Kittenplon, Inbal Lavi, Sharon Fogel +3
Text spotting end-to-end methods have recently gained attention in the literature due to the benefits of jointly optimizing the text detection and recognition components. Existing…
DocFormer: End-to-End Transformer for Document Understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota +2
We present DocFormer -- a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand docu…
On Calibration of Scene-Text Recognition Models
Ron Slossberg, Oron Anschel, Amir Markovitz +6
In this work, we study the problem of word-level confidence calibration for scene-text recognition (STR). Although the topic of confidence calibration has been an active research a…
Sequence-to-Sequence Contrastive Learning for Text Recognition
Aviad Aberdam, Ron Litman, Shahar Tsiper +5
We propose a framework for sequence-to-sequence contrastive learning (SeqCLR) of visual representations, which we apply to text recognition. To account for the sequence-to-sequence…
A Comprehensive Study of Deep Video Action Recognition
Yi Zhu, Xinyu Li, Chunhui Liu +7
Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks t…