Publications (5)
StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training
Yuechen Yu, Yulin Li, Chengquan Zhang +7
In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-tr…
GridFormer: Towards Accurate Table Structure Recognition via Grid Prediction
Pengyuan Lyu, Weihong Ma, Hongyi Wang +5
All tables can be represented as grids. Based on this observation, we propose GridFormer, a novel approach for interpreting unconstrained table structures by predicting the vertex…
TRUST: An Accurate and End-to-End Table structure Recognizer Using Splitting-based Transformers
Zengyuan Guo, Yuechen Yu, Pengyuan Lv +6
Table structure recognition is a crucial part of document image analysis domain. Its difficulty lies in the need to parse the physical coordinates and logical indices of each cell…
ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao +24
Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…
Deformable Siamese Attention Networks for Visual Object Tracking
Yuechen Yu, Yilei Xiong, Weilin Huang +1
Siamese-based trackers have achieved excellent performance on visual object tracking. However, the target template is not updated online, and the features of the target template an…