23 citations · 26 across the 2 of their papers we have counts for
12 papers · 1 filter
Emu: Generative Pretraining in Multimodality
Quan Sun, Qiying Yu, Yufeng Cui +7
We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-mo…
Fine-Grained Visual Prompting
Lingfeng Yang, Yueze Wang, Xiang Li +2
Vision-Language Models (VLMs), such as CLIP, have demonstrated impressive zero-shot transfer capabilities in image-level visual perception. However, these models have shown limited…
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
Yang Liu, Muzhi Zhu, Hengtao Li +3
Powered by large-scale pre-training, vision foundation models exhibit significant potential in open-world image understanding. However, unlike large language models that excel at d…
SegGPT: Segmenting Everything In Context
Xinlong Wang, Xiaosong Zhang, Yue Cao +3
We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates di…
EVA-CLIP: Improved Training Techniques for CLIP at Scale
Quan Sun, Yuxin Fang, Ledell Wu +2
Contrastive language-image pre-training, CLIP for short, has gained increasing attention for its potential in various scenarios. In this paper, we propose EVA-CLIP, a series of mod…
EVA-02: A Visual Representation for Neon Genesis
Yuxin Fang, Quan Sun, Xinggang Wang +3
We launch EVA-02, a next-generation Transformer-based visual representation pre-trained to reconstruct strong and robust language-aligned vision features via masked image modeling.…