activity
20192023
most citedEVA: Exploring the Limits of Masked Visual Representation Learning at Scale

23 citations · 26 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2023

Emu: Generative Pretraining in Multimodality

Quan Sun, Qiying Yu, Yufeng Cui +7

We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-mo…

cs.CV2023

Fine-Grained Visual Prompting

Lingfeng Yang, Yueze Wang, Xiang Li +2

Vision-Language Models (VLMs), such as CLIP, have demonstrated impressive zero-shot transfer capabilities in image-level visual perception. However, these models have shown limited…

cs.CV2023

Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching

Yang Liu, Muzhi Zhu, Hengtao Li +3

Powered by large-scale pre-training, vision foundation models exhibit significant potential in open-world image understanding. However, unlike large language models that excel at d…

cs.CV202362 cited

SegGPT: Segmenting Everything In Context

Xinlong Wang, Xiaosong Zhang, Yue Cao +3

We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates di…

cs.CV202381 cited

EVA-CLIP: Improved Training Techniques for CLIP at Scale

Quan Sun, Yuxin Fang, Ledell Wu +2

Contrastive language-image pre-training, CLIP for short, has gained increasing attention for its potential in various scenarios. In this paper, we propose EVA-CLIP, a series of mod…

cs.CV2023

EVA-02: A Visual Representation for Neon Genesis

Yuxin Fang, Quan Sun, Xinggang Wang +3

We launch EVA-02, a next-generation Transformer-based visual representation pre-trained to reconstruct strong and robust language-aligned vision features via masked image modeling.…