18 citations · 33 across the 8 of their papers we have counts for
8 papers · 1 filter
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
Guosheng Zhang, Keyao Wang, Haixiao Yue +5
Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tas…
GridFormer: Towards Accurate Table Structure Recognition via Grid Prediction
Pengyuan Lyu, Weihong Ma, Hongyi Wang +5
All tables can be represented as grids. Based on this observation, we propose GridFormer, a novel approach for interpreting unconstrained table structures by predicting the vertex…
Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation
Huan Liu, Qiang Chen, Zichang Tan +9
In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.…
Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation Learning
Xugong Qin, Pengyuan Lyu, Chengquan Zhang +5
Due to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection…
ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao +24
Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, an…
Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding
Mingliang Zhai, Yulin Li, Xiameng Qin +6
Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the…