activity
20212024
most citedStrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

18 citations · 22 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2024

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Hannan Lu, Xiaohe Wu, Shudong Wang +5

Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Exist…

cs.CV2024

Collaborative Position Reasoning Network for Referring Image Segmentation

Jianjian Cao, Beiya Dai, Yulin Li +2

Given an image and a natural language expression as input, the goal of referring image segmentation is to segment the foreground masks of the entities referred by the expression. E…

cs.CV20231 cited

MataDoc: Margin and Text Aware Document Dewarping for Arbitrary Boundary

Beiya Dai, Xing li, Qunyi Xie +5

Document dewarping from a distorted camera-captured image is of great value for OCR and document understanding. The document boundary plays an important role which is more evident…

cs.CV2023

Fast-StrucTexT: An Efficient Hourglass Transformer with Modality-guided Dynamic Token Merge for Document Understanding

Mingliang Zhai, Yulin Li, Xiameng Qin +6

Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the…

cs.CV202318 cited

StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

Yuechen Yu, Yulin Li, Chengquan Zhang +7

In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-tr…

cs.CV20213 cited

Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question Answering

JianJian Cao, Xiameng Qin, Sanyuan Zhao +1

Answering semantically-complicated questions according to an image is challenging in Visual Question Answering (VQA) task. Although the image can be well represented by deep learni…