most citedVisual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

189 citations · 247 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV20232 cited

LayoutNUWA: Revealing the Hidden Layout Expertise of Large Language Models

Zecheng Tang, Chenfei Wu, Juntao Li +1

Graphic layout generation, a growing research field, plays a significant role in user engagement and information perception. Existing methods primarily treat layout generation as a…

cs.CV2023

ORES: Open-vocabulary Responsible Visual Synthesis

Minheng Ni, Chenfei Wu, Xiaodong Wang +4

Avoiding synthesizing specific visual concepts is an essential challenge in responsible visual synthesis. However, the visual concept that needs to be avoided for responsible visua…

cs.CV20239 cited

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

Shengming Yin, Chenfei Wu, Jian Liang +4

Controllable video generation has gained significant attention in recent years. However, two main limitations persist: Firstly, most existing works focus on either text, image, or…

cs.CV2023

ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning

Xiao Xu, Bei Li, Chenfei Wu +6

Two-Tower Vision-Language (VL) models have shown promising improvements on various downstream VL tasks. Although the most advanced work improves performance by building bridges bet…

cs.CV20231 cited

NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

Shengming Yin, Chenfei Wu, Huan Yang +13

In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment seq…

cs.CV2023189 cited

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Chenfei Wu, Shengming Yin, Weizhen Qi +3

ChatGPT is attracting a cross-field interest as it provides a language interface with remarkable conversational competency and reasoning capabilities across many domains. However,…