most citedRegionCLIP: Region-based Language-Image Pretraining

13 citations · 34 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV202310 cited

GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

An Yan, Zhengyuan Yang, Wanrong Zhu +9

We present MM-Navigator, a GPT-4V-based agent for the smartphone graphical user interface (GUI) navigation task. MM-Navigator can interact with a smartphone screen as human users,…

cs.CV20239 cited

Robust and Interpretable Medical Image Classifiers via Concept Bottleneck Models

An Yan, Yu Wang, Yiwu Zhong +8

Medical image classification is a critical problem for healthcare, with the potential to alleviate the workload of doctors and facilitate diagnoses of patients. However, two challe…

cs.CV20231 cited

Learning Concise and Descriptive Attributes for Visual Recognition

An Yan, Yu Wang, Yiwu Zhong +6

Recent advances in foundation models present new opportunities for interpretable visual recognition -- one can first query Large Language Models (LLMs) to obtain a set of attribute…

cs.CV20231 cited

Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations

Yiwu Zhong, Licheng Yu, Yang Bai +3

The abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn vi…

cs.CV202113 cited

RegionCLIP: Region-based Language-Image Pretraining

Yiwu Zhong, Jianwei Yang, Pengchuan Zhang +8

Contrastive language-image pretraining (CLIP) using image-text pairs has achieved impressive results on image classification in both zero-shot and transfer learning settings. Howev…