14 citations · 49 across the 22 of their papers we have counts for
26 papers · 1 filter
Turning a CLIP Model into a Scene Text Spotter
Wenwen Yu, Yuliang Liu, Xingkui Zhu +3
We exploit the potential of the large-scale Contrastive Language-Image Pretraining (CLIP) model to enhance scene text detection and spotting tasks, transforming it into a robust ba…
ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer
Mingxin Huang, Jiaxin Zhang, Dezhi Peng +5
In recent years, end-to-end scene text spotting approaches are evolving to the Transformer-based framework. While previous studies have shown the crucial importance of the intrinsi…
Visual Information Extraction in the Wild: Practical Dataset and End-to-end Solution
Jianfeng Kuang, Wei Hua, Dingkang Liang +4
Visual information extraction (VIE), which aims to simultaneously perform OCR and information extraction in a unified framework, has drawn increasing attention due to its essential…
Looking and Listening: Audio Guided Text Recognition
Wenwen Yu, Mingyu Liu, Biao Yang +5
Text recognition in the wild is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest vision and language processing are effective…
ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao +24
Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…
SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
Dingyuan Zhang, Dingkang Liang, Hongcheng Yang +4
With the development of large language models, many remarkable linguistic systems like ChatGPT have thrived and achieved astonishing success on many tasks, showing the incredible p…