16 citations · 46 across the 4 of their papers we have counts for
7 papers
MMOCR: A Comprehensive Toolbox for Text Detection, Recognition and Understanding
Zhanghui Kuang, Hongbin Sun, Zhizhong Li +10
We present MMOCR-an open-source toolbox which provides a comprehensive pipeline for text detection and recognition, as well as their downstream tasks such as named entity recogniti…
Vision Transformer with Progressive Sampling
Xiaoyu Yue, Shuyang Sun, Zhanghui Kuang +4
Transformers with powerful global relation modeling abilities have been introduced to fundamental computer vision tasks recently. As a typical example, the Vision Transformer (ViT)…
Spatial Dual-Modality Graph Reasoning for Key Information Extraction
Hongbin Sun, Zhanghui Kuang, Xiaoyu Yue +2
Key information extraction from document images is of paramount importance in office automation. Conventional template matching based approaches fail to generalize well to document…
HOSE-Net: Higher Order Structure Embedded Network for Scene Graph Generation
Meng Wei, Chun Yuan, Xiaoyu Yue +1
Scene graph generation aims to produce structured representations for images, which requires to understand the relations between objects. Due to the continuous nature of deep neura…
RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition
Xiaoyu Yue, Zhanghui Kuang, Chenhao Lin +2
The attention-based encoder-decoder framework has recently achieved impressive results for scene text recognition, and many variants have emerged with improvements in recognition q…
Geometry Normalization Networks for Accurate Scene Text Detection
Youjiang Xu, Jiaqi Duan, Zhanghui Kuang +4
Large geometry (e.g., orientation) variances are the key challenges in the scene text detection. In this work, we first conduct experiments to investigate the capacity of networks…