activity
20182021
most citedDeVLBert: Learning Deconfounded Visio-Linguistic Representations

64 citations · 225 across the 19 of their papers we have counts for

collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2021

VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations

Peng Zhang, Can Li, Liang Qiao +4

Document layout analysis is crucial for understanding document structures. On this task, vision and semantics of documents, and relations between layout components contribute to th…

cs.CV2021

Modeling High-order Interactions across Multi-interests for Micro-video Recommendation

Dong Yao, Shengyu Zhang, Zhou Zhao +4

Personalized recommendation system has become pervasive in various video platform. Many effective methods have been proposed, but most of them didn't capture the user's multi-level…

cs.CV2021

Reciprocal Feature Learning via Explicit and Implicit Tasks in Scene Text Recognition

Hui Jiang, Yunlu Xu, Zhanzhan Cheng +5

Text recognition is a popular topic for its broad applications. In this work, we excavate the implicit task, character counting within the traditional text recognition, without add…

cs.CV2020

MANGO: A Mask Attention Guided One-Stage Scene Text Spotter

Liang Qiao, Ying Chen, Zhanzhan Cheng +4

Recently end-to-end scene text spotting has become a popular research topic due to its advantages of global optimization and high maintainability in real applications. Most methods…

cs.CV2020

MGD-GAN: Text-to-Pedestrian generation through Multi-Grained Discrimination

Shengyu Zhang, Donghui Wang, Zhou Zhao +3

In this paper, we investigate the problem of text-to-pedestrian synthesis, which has many potential applications in art, design, and video surveillance. Existing methods for text-t…

cs.CV202064 cited

DeVLBert: Learning Deconfounded Visio-Linguistic Representations

Shengyu Zhang, Tan Jiang, Tan Wang +6

In this paper, we propose to investigate the problem of out-of-domain visio-linguistic pretraining, where the pretraining data distribution differs from that of downstream data on…