activity
20192026
most citedVisual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

189 citations · 310 across the 30 of their papers we have counts for

collaborators
Showing 2022 · cs.CVShow all

7 papers · 2 filters

cs.CV2022

ReCo: Region-Controlled Text-to-Image Generation

Zhengyuan Yang, Jianfeng Wang, Zhe Gan +8

Recently, large-scale text-to-image (T2I) models have shown impressive performance in generating high-fidelity images, but with limited controllability, e.g., precisely specifying…

cs.CV2022★ 1 cited

HORIZON: High-Resolution Semantically Controlled Panorama Synthesis

Kun Yan, Lei Ji, Chenfei Wu +4

Panorama synthesis endeavors to craft captivating 360-degree visual landscapes, immersing users in the heart of virtual worlds. Nevertheless, contemporary panoramic synthesis techn…

cs.CV2022★ 25 cited

NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual Synthesis

Chenfei Wu, Jian Liang, Xiaowei Hu +6

In this paper, we present NUWA-Infinity, a generative model for infinite visual synthesis, which is defined as the task of generating arbitrarily-sized high-resolution images or lo…

cs.CV2022★ 3 cited

BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning

Xiao Xu, Chenfei Wu, Shachar Rosenman +3

Vision-Language (VL) models with the Two-Tower architecture have dominated visual-language representation learning in recent years. Current VL models either use lightweight uni-mod…

cs.CV2022★ 11 cited

DiVAE: Photorealistic Images Synthesis with Denoising Diffusion Decoder

Jie Shi, Chenfei Wu, Jian Liang +2

Recently most successful image synthesis models are multi stage process to combine the advantages of different methods, which always includes a VAE-like model for faithfully recons…

cs.CV2022★ 3 cited

VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers

Estelle Aflalo, Meng Du, Shao-Yen Tseng +4

Breakthroughs in transformer-based models have revolutionized not only the NLP field, but also vision and multimodal systems. However, although visualization and interpretability t…