6 citations · 14 across the 5 of their papers we have counts for
4 papers · 1 filter
DocTr: Document Transformer for Structured Information Extraction in Documents
Haofu Liao, Aruni RoyChowdhury, Weijian Li +6
We present a new formulation for structured information extraction (SIE) from visually rich documents. It aims to address the limitations of existing IOB tagging or graph-based for…
PolyFormer: Referring Image Segmentation as Sequential Polygon Generation
Jiang Liu, Hui Ding, Zhaowei Cai +4
In this work, instead of directly predicting the pixel-level segmentation masks, the problem of referring image segmentation is formulated as sequential polygon generation, and the…
MATrIX -- Modality-Aware Transformer for Information eXtraction
Thomas Delteil, Edouard Belval, Lei Chen +2
We present MATrIX - a Modality-Aware Transformer for Information eXtraction in the Visual Document Understanding (VDU) domain. VDU covers information extraction from visually rich…
Visual Relationship Detection Using Part-and-Sum Transformers with Composite Queries
Qi Dong, Zhuowen Tu, Haofu Liao +3
Computer vision applications such as visual relationship detection and human object interaction can be formulated as a composite (structured) set detection problem in which both th…