1 citations · 1 across the 4 of their papers we have counts for
4 papers
StrucTexTv3: An Efficient Vision-Language Model for Text-rich Image Perception, Comprehension, and Beyond
Pengyuan Lyu, Yulin Li, Hao Zhou +8
Text-rich images have significant and extensive value, deeply integrated into various aspects of human life. Notably, both visual cues and linguistic symbols in text-rich images pl…
GridFormer: Towards Accurate Table Structure Recognition via Grid Prediction
Pengyuan Lyu, Weihong Ma, Hongyi Wang +5
All tables can be represented as grids. Based on this observation, we propose GridFormer, a novel approach for interpreting unconstrained table structures by predicting the vertex…
Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation Learning
Xugong Qin, Pengyuan Lyu, Chengquan Zhang +5
Due to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection…
ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao +24
Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…