activity
20102023
most citedStrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

18 citations · 33 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025

Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models

Guosheng Zhang, Keyao Wang, Haixiao Yue +5

Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tas…

cs.CV2023

GridFormer: Towards Accurate Table Structure Recognition via Grid Prediction

Pengyuan Lyu, Weihong Ma, Hongyi Wang +5

All tables can be represented as grids. Based on this observation, we propose GridFormer, a novel approach for interpreting unconstrained table structures by predicting the vertex…

cs.CV20233 cited

Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation

Huan Liu, Qiang Chen, Zichang Tan +9

In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.…

cs.CV20231 cited

Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation Learning

Xugong Qin, Pengyuan Lyu, Chengquan Zhang +5

Due to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection…

cs.CV2023

ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images

Wenwen Yu, Chengquan Zhang, Haoyu Cao +24

Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…

cs.CV2023

Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding

Mingliang Zhai, Yulin Li, Xiameng Qin +6

Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the…