80 citations · 116 across the 7 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023★ 80 cited
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
Zhengyuan Yang, Linjie Li, Jianfeng Wang +7
We propose MM-REACT, a system paradigm that integrates ChatGPT with a pool of vision experts to achieve multimodal reasoning and action. In this paper, we define and explore a comp…
cs.CV2023★ 18 cited
VA-DepthNet: A Variational Approach to Single Image Depth Prediction
Ce Liu, Suryansh Kumar, Shuhang Gu +2
We introduce VA-DepthNet, a simple, effective, and accurate deep neural network approach for the single-image depth prediction (SIDP) problem. The proposed approach advocates using…
cs.CV2023
Learning Customized Visual Models with Retrieval-Augmented Knowledge
Haotian Liu, Kilho Son, Jianwei Yang +4
Image-text contrastive learning models such as CLIP have demonstrated strong task transfer ability. The high generality and usability of these visual models is achieved via a web-s…