activity
20182023
most citedFocal Self-attention for Local-Global Interactions in Vision Transformers

268 citations · 2.1k across the 78 of their papers we have counts for

collaborators
Showing cs.CVShow all

25 papers · 1 filter

cs.CV2023★ 51 cited

Semantic-SAM: Segment and Recognize Anything at Any Granularity

Feng Li, Hao Zhang, Peize Sun +6

In this paper, we introduce Semantic-SAM, a universal image segmentation model to enable segment and recognize anything at any desired granularity. Our model offers two key advanta…

cs.CV2023★ 231 cited

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Chunyuan Li, Cliff Wong, Sheng Zhang +6

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversation…

cs.CV2023★ 2 cited

ArK: Augmented Reality with Knowledge Interactive Emergent Ability

Qiuyuan Huang, Jae Sung Park, Abhinav Gupta +8

Despite the growing adoption of mixed reality and interactive AI agents, it remains challenging for these systems to generate high quality 2D/3D scenes in unseen environments. The…

cs.CV2023★ 154 cited

Segment Everything Everywhere All at Once

Xueyan Zou, Jianwei Yang, Hao Zhang +6

In this work, we present SEEM, a promptable and interactive model for segmenting everything everywhere all at once in an image, as shown in Fig.1. In SEEM, we propose a novel decod…

cs.CV2023★ 2 cited

A Simple Framework for Open-Vocabulary Segmentation and Detection

Hao Zhang, Feng Li, Xueyan Zou +5

We present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets. To bridge the gap of voca…

cs.CV2023

Learning Customized Visual Models with Retrieval-Augmented Knowledge

Haotian Liu, Kilho Son, Jianwei Yang +4

Image-text contrastive learning models such as CLIP have demonstrated strong task transfer ability. The high generality and usability of these visual models is achieved via a web-s…