most citedOFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

258 citations · 321 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2022★ 3 cited

Single Stage Virtual Try-on via Deformable Attention Flows

Shuai Bai, Huiling Zhou, Zhikang Li +2

Virtual try-on aims to generate a photo-realistic fitting result given an in-shop garment and a reference person image. Existing methods usually build up multi-stage frameworks to…

cs.CV2022★ 3 cited

M6-Fashion: High-Fidelity Multi-modal Image Generation and Editing

Zhikang Li, Huiling Zhou, Shuai Bai +3

The fashion industry has diverse applications in multi-modal image generation and editing. It aims to create a desired high-fidelity image with the multi-modal conditional signal a…

cs.CV2022★ 258 cited

OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Peng Wang, An Yang, Rui Men +7

In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Mo…

cs.IR2021★ 1 cited

Cross-domain User Preference Learning for Cold-start Recommendation

Huiling Zhou, Jie Liu, Zhikang Li +2

Cross-domain cold-start recommendation is an increasingly emerging issue for recommender systems. Existing works mainly focus on solving either cross-domain user recommendation or…

cs.CV2021★ 8 cited

M6-UFC: Unifying Multi-Modal Controls for Conditional Image Synthesis via Non-Autoregressive Generative Transformers

Zhu Zhang, Jianxin Ma, Chang Zhou +6

Conditional image synthesis aims to create an image according to some multi-modal guidance in the forms of textual descriptions, reference images, and image blocks to preserve, as…

cs.CL2021★ 48 cited

M6: A Chinese Multimodal Pretrainer

Junyang Lin, Rui Men, An Yang +22

In this work, we construct the largest dataset for multimodal pretraining in Chinese, which consists of over 1.9TB images and 292GB texts that cover a wide range of domains. We pro…