activity
20152024
most citedHigh Performance Offline Handwritten Chinese Character Recognition Using GoogLeNet and Directional Feature Maps

21 citations · 28 across the 8 of their papers we have counts for

collaborators

12 papers

cs.CV2024

Dynamic Relation Transformer for Contextual Text Block Detection

Jiawei Wang, Shunchi Zhang, Kai Hu +4

Contextual Text Block Detection (CTBD) is the task of identifying coherent text blocks within the complexity of natural scenes. Previous methodologies have treated CTBD as either a…

cs.CL2024

UniVIE: A Unified Label Space Approach to Visual Information Extraction from Form-like Documents

Kai Hu, Jiawei Wang, Weihong Lin +3

Existing methods for Visual Information Extraction (VIE) from form-like documents typically fragment the process into separate subtasks, such as key information extraction, key-val…

cs.CV2024★ 2 cited

Detect-Order-Construct: A Tree Construction based Approach for Hierarchical Document Structure Analysis

Jiawei Wang, Kai Hu, Zhuoyao Zhong +2

Document structure analysis (aka document layout analysis) is crucial for understanding the physical layout and logical structure of documents, with applications in information ret…

cs.CV2023★ 1 cited

Exploring Predicate Visual Context in Detecting Human-Object Interactions

Frederic Z. Zhang, Yuhui Yuan, Dylan Campbell +2

Recently, the DETR framework has emerged as the dominant approach for human--object interaction (HOI) research. In particular, two-stage transformer-based HOI detectors are amongst…

cs.CL2023★ 1 cited

A Question-Answering Approach to Key Value Pair Extraction from Form-like Document Images

Kai Hu, Zhuoyuan Wu, Zhuoyao Zhong +3

In this paper, we present a new question-answering (QA) based key-value pair extraction approach, called KVPFormer, to robustly extracting key-value relationships between entities…

cs.CL2021★ 3 cited

ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Information Extraction from Documents

Weihong Lin, Qifang Gao, Lei Sun +4

Recent grid-based document representations like BERTgrid allow the simultaneous encoding of the textual and layout information of a document in a 2D feature map so that state-of-th…