241 citations · 399 across the 15 of their papers we have counts for
19 papers · 1 filter
Emu: Generative Pretraining in Multimodality
Quan Sun, Qiying Yu, Yufeng Cui +7
We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-mo…
SVIT: Scaling up Visual Instruction Tuning
Bo Zhao, Boya Wu, Muyang He +1
Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Al…
Pushing the Limits of 3D Shape Generation at Scale
Yu Wang, Xuelin Qian, Jingyang Huo +3
We present a significant breakthrough in 3D shape generation by scaling it to unprecedented dimensions. Through the adaptation of the Auto-Regressive model and the utilization of l…
SegGPT: Segmenting Everything In Context
Xinlong Wang, Xiaosong Zhang, Yue Cao +3
We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates di…
EVA-02: A Visual Representation for Neon Genesis
Yuxin Fang, Quan Sun, Xinggang Wang +3
We launch EVA-02, a next-generation Transformer-based visual representation pre-trained to reconstruct strong and robust language-aligned vision features via masked image modeling.…
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
Yuxin Fang, Wen Wang, Binhui Xie +6
We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to re…