81 citations · 81 across the 1 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023
Emu: Generative Pretraining in Multimodality
Quan Sun, Qiying Yu, Yufeng Cui +7
We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-mo…
cs.CV2023★ 81 cited
EVA-CLIP: Improved Training Techniques for CLIP at Scale
Quan Sun, Yuxin Fang, Ledell Wu +2
Contrastive language-image pre-training, CLIP for short, has gained increasing attention for its potential in various scenarios. In this paper, we propose EVA-CLIP, a series of mod…
cs.CV2023
EVA-02: A Visual Representation for Neon Genesis
Yuxin Fang, Quan Sun, Xinggang Wang +3
We launch EVA-02, a next-generation Transformer-based visual representation pre-trained to reconstruct strong and robust language-aligned vision features via masked image modeling.…