787 citations · 1.8k across the 39 of their papers we have counts for
7 papers · 1 filter
M6-Fashion: High-Fidelity Multi-modal Image Generation and Editing
Zhikang Li, Huiling Zhou, Shuai Bai +3
The fashion industry has diverse applications in multi-modal image generation and editing. It aims to create a desired high-fidelity image with the multi-modal conditional signal a…
Connecting Language and Vision for Natural Language-Based Vehicle Retrieval
Shuai Bai, Zhedong Zheng, Xiaohan Wang +5
Vehicle search is one basic task for the efficient traffic management in terms of the AI City. Most existing practices focus on the image-based vehicle matching, including vehicle…
CogView: Mastering Text-to-Image Generation via Transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong +8
Text-to-Image generation in the general domain has long been an open problem, which requires both a powerful generative model and cross-modal understanding. We propose CogView, a 4…
DeVLBert: Learning Deconfounded Visio-Linguistic Representations
Shengyu Zhang, Tan Jiang, Tan Wang +6
In this paper, we propose to investigate the problem of out-of-domain visio-linguistic pretraining, where the pretraining data distribution differs from that of downstream data on…
Poet: Product-oriented Video Captioner for E-commerce
Shengyu Zhang, Ziqi Tan, Jin Yu +6
In e-commerce, a growing number of user-generated videos are used for product promotion. How to generate video descriptions that narrate the user-preferred product characteristics…
Comprehensive Information Integration Modeling Framework for Video Titling
Shengyu Zhang, Ziqi Tan, Jin Yu +6
In e-commerce, consumer-generated videos, which in general deliver consumers' individual preferences for the different aspects of certain products, are massive in volume. To recomm…