activity
20142023
most citedMM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

80 citations · 93 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20231 cited

NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

Shengming Yin, Chenfei Wu, Huan Yang +13

In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment seq…

cs.CV202380 cited

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Zhengyuan Yang, Linjie Li, Jianfeng Wang +7

We propose MM-REACT, a system paradigm that integrates ChatGPT with a pool of vision experts to achieve multimodal reasoning and action. In this paper, we define and explore a comp…

cs.CV20212 cited

The Overlooked Classifier in Human-Object Interaction Recognition

Ying Jin, Yinpeng Chen, Lijuan Wang +5

Human-Object Interaction (HOI) recognition is challenging due to two factors: (1) significant imbalance across classes and (2) requiring multiple labels per image. This paper shows…

cs.CV20218 cited

Injecting Semantic Concepts into End-to-End Image Captioning

Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu +5

Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Re…

cs.CV20142 cited

Optimized Cartesian -Means

Jianfeng Wang, Jingdong Wang, Jingkuan Song +3

Product quantization-based approaches are effective to encode high-dimensional data points for approximate nearest neighbor search. The space is decomposed into a Cartesian product…