most citedDeVLBert: Learning Deconfounded Visio-Linguistic Representations

64 citations · 115 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV202064 cited

DeVLBert: Learning Deconfounded Visio-Linguistic Representations

Shengyu Zhang, Tan Jiang, Tan Wang +6

In this paper, we propose to investigate the problem of out-of-domain visio-linguistic pretraining, where the pretraining data distribution differs from that of downstream data on…

cs.CV202025 cited

Poet: Product-oriented Video Captioner for E-commerce

Shengyu Zhang, Ziqi Tan, Jin Yu +6

In e-commerce, a growing number of user-generated videos are used for product promotion. How to generate video descriptions that narrate the user-preferred product characteristics…

cs.SI2020

A Multi-Semantic Metapath Model for Large Scale Heterogeneous Network Representation Learning

Xuandong Zhao, Jinbao Xue, Jin Yu +2

Network Embedding has been widely studied to model and manage data in a variety of real-world applications. However, most existing works focus on networks with single-typed nodes o…

cs.CV202024 cited

Comprehensive Information Integration Modeling Framework for Video Titling

Shengyu Zhang, Ziqi Tan, Jin Yu +6

In e-commerce, consumer-generated videos, which in general deliver consumers' individual preferences for the different aspects of certain products, are massive in volume. To recomm…

cs.CV20202 cited

Grounded and Controllable Image Completion by Incorporating Lexical Semantics

Shengyu Zhang, Tan Jiang, Qinghao Huang +7

In this paper, we present an approach, namely Lexical Semantic Image Completion (LSIC), that may have potential applications in art, design, and heritage conservation, among severa…