activity
20212024
most citedMM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

80 citations · 153 across the 16 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV202316 cited

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Bin Xiao, Haiping Wu, Weijian Xu +6

We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for a variety of computer vision and vision-language tasks. While existing larg…

cs.CV202329 cited

Any-to-Any Generation via Composable Diffusion

Zineng Tang, Ziyi Yang, Chenguang Zhu +2

We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any comb…

cs.CV202380 cited

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Zhengyuan Yang, Linjie Li, Jianfeng Wang +7

We propose MM-REACT, a system paradigm that integrates ChatGPT with a pool of vision experts to achieve multimodal reasoning and action. In this paper, we define and explore a comp…

cs.CV20228 cited

CLIP-Event: Connecting Text and Images with Event Structures

Manling Li, Ruochen Xu, Shuohang Wang +6

Vision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text. While existing v…

cs.CV20215 cited

MLP Architectures for Vision-and-Language Modeling: An Empirical Study

Yixin Nie, Linjie Li, Zhe Gan +6

We initiate the first empirical study on the use of MLP architectures for vision-and-language (VL) fusion. Through extensive experiments on 5 VL tasks and 5 robust VQA benchmarks,…