64 citations · 225 across the 19 of their papers we have counts for
14 papers · 1 filter
VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations
Peng Zhang, Can Li, Liang Qiao +4
Document layout analysis is crucial for understanding document structures. On this task, vision and semantics of documents, and relations between layout components contribute to th…
Modeling High-order Interactions across Multi-interests for Micro-video Recommendation
Dong Yao, Shengyu Zhang, Zhou Zhao +4
Personalized recommendation system has become pervasive in various video platform. Many effective methods have been proposed, but most of them didn't capture the user's multi-level…
Reciprocal Feature Learning via Explicit and Implicit Tasks in Scene Text Recognition
Hui Jiang, Yunlu Xu, Zhanzhan Cheng +5
Text recognition is a popular topic for its broad applications. In this work, we excavate the implicit task, character counting within the traditional text recognition, without add…
MANGO: A Mask Attention Guided One-Stage Scene Text Spotter
Liang Qiao, Ying Chen, Zhanzhan Cheng +4
Recently end-to-end scene text spotting has become a popular research topic due to its advantages of global optimization and high maintainability in real applications. Most methods…
MGD-GAN: Text-to-Pedestrian generation through Multi-Grained Discrimination
Shengyu Zhang, Donghui Wang, Zhou Zhao +3
In this paper, we investigate the problem of text-to-pedestrian synthesis, which has many potential applications in art, design, and video surveillance. Existing methods for text-t…
DeVLBert: Learning Deconfounded Visio-Linguistic Representations
Shengyu Zhang, Tan Jiang, Tan Wang +6
In this paper, we propose to investigate the problem of out-of-domain visio-linguistic pretraining, where the pretraining data distribution differs from that of downstream data on…