25 citations · 33 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 6 cited
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Shilong Liu, Hao Cheng, Haotian Liu +10
LLaVA-Plus is a general-purpose multimodal assistant that expands the capabilities of large multimodal models. It maintains a skill repository of pre-trained vision and vision-lang…
cs.CV2023★ 25 cited
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Jianwei Yang, Hao Zhang, Feng Li +3
We present Set-of-Mark (SoM), a new visual prompting method, to unleash the visual grounding abilities of large multimodal models (LMMs), such as GPT-4V. As illustrated in Fig. 1 (…
cs.CV2023★ 2 cited
A Simple Framework for Open-Vocabulary Segmentation and Detection
Hao Zhang, Feng Li, Xueyan Zou +5
We present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets. To bridge the gap of voca…