80 citations · 137 across the 11 of their papers we have counts for
11 papers
StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis
Zecheng Tang, Chenfei Wu, Zekai Zhang +8
To leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model…
GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
An Yan, Zhengyuan Yang, Wanrong Zhu +9
We present MM-Navigator, a GPT-4V-based agent for the smartphone graphical user interface (GUI) navigation task. MM-Navigator can interact with a smartphone screen as human users,…
OpenLEAF: Open-Domain Interleaved Image-Text Generation and Evaluation
Jie An, Zhengyuan Yang, Linjie Li +5
This work investigates a challenging task named open-domain interleaved image-text generation, which generates interleaved texts and images following an input query. We propose a n…
MM-VID: Advancing Video Understanding with GPT-4V(ision)
Kevin Lin, Faisal Ahmed, Linjie Li +9
We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video unders…
Adaptive Human Matting for Dynamic Videos
Chung-Ching Lin, Jiang Wang, Kun Luo +4
The most recent efforts in video matting have focused on eliminating trimap dependency since trimap annotations are expensive and trimap-based methods are less adaptable for real-t…
NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation
Shengming Yin, Chenfei Wu, Huan Yang +13
In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment seq…