activity
20212024
most citedMM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

80 citations · 137 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CV2024

StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis

Zecheng Tang, Chenfei Wu, Zekai Zhang +8

To leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model…

cs.CV202310 cited

GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

An Yan, Zhengyuan Yang, Wanrong Zhu +9

We present MM-Navigator, a GPT-4V-based agent for the smartphone graphical user interface (GUI) navigation task. MM-Navigator can interact with a smartphone screen as human users,…

cs.CV20231 cited

OpenLEAF: Open-Domain Interleaved Image-Text Generation and Evaluation

Jie An, Zhengyuan Yang, Linjie Li +5

This work investigates a challenging task named open-domain interleaved image-text generation, which generates interleaved texts and images following an input query. We propose a n…

cs.CV202310 cited

MM-VID: Advancing Video Understanding with GPT-4V(ision)

Kevin Lin, Faisal Ahmed, Linjie Li +9

We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video unders…

cs.CV2023

Adaptive Human Matting for Dynamic Videos

Chung-Ching Lin, Jiang Wang, Kun Luo +4

The most recent efforts in video matting have focused on eliminating trimap dependency since trimap annotations are expensive and trimap-based methods are less adaptable for real-t…

cs.CV20231 cited

NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

Shengming Yin, Chenfei Wu, Huan Yang +13

In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment seq…