most citedVisual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

189 citations · 191 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20241 cited

Using Left and Right Brains Together: Towards Vision and Language Planning

Jun Cen, Chenfei Wu, Xiao Liu +6

Large Language Models (LLMs) and Large Multi-modality Models (LMMs) have demonstrated remarkable decision masking capabilities on a variety of tasks. However, they inherently opera…

cs.CV2024

StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis

Zecheng Tang, Chenfei Wu, Zekai Zhang +8

To leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model…

cs.CV2023

ORES: Open-vocabulary Responsible Visual Synthesis

Minheng Ni, Chenfei Wu, Xiaodong Wang +4

Avoiding synthesizing specific visual concepts is an essential challenge in responsible visual synthesis. However, the visual concept that needs to be avoided for responsible visua…

cs.CV20231 cited

NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

Shengming Yin, Chenfei Wu, Huan Yang +13

In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment seq…

cs.CV2023189 cited

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Chenfei Wu, Shengming Yin, Weizhen Qi +3

ChatGPT is attracting a cross-field interest as it provides a language interface with remarkable conversational competency and reasoning capabilities across many domains. However,…

cs.CV2023

Learning 3D Photography Videos via Self-supervised Diffusion on Single Images

Xiaodong Wang, Chenfei Wu, Shengming Yin +9

3D photography renders a static image into a video with appealing 3D visual effects. Existing approaches typically first conduct monocular depth estimation, then render the input f…