activity
20222024
most citedStrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

18 citations · 23 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2024

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Hannan Lu, Xiaohe Wu, Shudong Wang +5

Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Exist…

cs.CV2024

Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection

Jiaming Li, Xiangru Lin, Wei Zhang +6

Current semi-supervised object detection (SSOD) algorithms typically assume class balanced datasets (PASCAL VOC etc.) or slightly class imbalanced datasets (MS-COCO, etc). This ass…

cs.CV20231 cited

Semi-DETR: Semi-Supervised Object Detection with Detection Transformers

Jiacheng Zhang, Xiangru Lin, Wei Zhang +6

We analyze the DETR-based framework on semi-supervised object detection (SSOD) and observe that (1) the one-to-one assignment strategy generates incorrect matching when the pseudo…

cs.CV20232 cited

Ambiguity-Resistant Semi-Supervised Learning for Dense Object Detection

Chang Liu, Weiming Zhang, Xiangru Lin +6

With basic Semi-Supervised Object Detection (SSOD) techniques, one-stage detectors generally obtain limited promotions compared with two-stage clusters. We experimentally find that…

cs.CV20232 cited

PSVT: End-to-End Multi-person 3D Pose and Shape Estimation with Progressive Video Transformers

Zhongwei Qiu, Yang Qiansheng, Jian Wang +6

Existing methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then per…

cs.CV202318 cited

StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

Yuechen Yu, Yulin Li, Chengquan Zhang +7

In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-tr…