2 citations · 3 across the 4 of their papers we have counts for
4 papers · 1 filter
ReCo: Region-Controlled Text-to-Image Generation
Zhengyuan Yang, Jianfeng Wang, Zhe Gan +8
Recently, large-scale text-to-image (T2I) models have shown impressive performance in generating high-fidelity images, but with limited controllability, e.g., precisely specifying…
NÜWA-LIP: Language Guided Image Inpainting with Defect-free VQGAN
Minheng Ni, Chenfei Wu, Haoyang Huang +3
Language guided image inpainting aims to fill in the defective regions of an image under the guidance of text while keeping non-defective regions unchanged. However, the encoding p…
GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
Chenfei Wu, Lun Huang, Qianxi Zhang +5
Generating videos from text is a challenging task due to its high computational requirements for training and infinite possible answers for evaluation. Existing works typically exp…
Deep Reason: A Strong Baseline for Real-World Visual Reasoning
Chenfei Wu, Yanzhao Zhou, Gen Li +3
This paper presents a strong baseline for real-world visual reasoning (GQA), which achieves 60.93% in GQA 2019 challenge and won the sixth place. GQA is a large dataset with 22M qu…