activity
20162023
most citedLearning Open Set Network with Discriminative Reciprocal Points

241 citations · 1.1k across the 35 of their papers we have counts for

collaborators
Showing 2023 · cs.CVShow all

6 papers · 2 filters

cs.CV2023★ 29 cited

Emu: Generative Pretraining in Multimodality

Quan Sun, Qiying Yu, Yufeng Cui +7

We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-mo…

cs.CV2023★ 14 cited

SVIT: Scaling up Visual Instruction Tuning

Bo Zhao, Boya Wu, Muyang He +1

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Al…

cs.CV2023★ 4 cited

Pushing the Limits of 3D Shape Generation at Scale

Yu Wang, Xuelin Qian, Jingyang Huo +3

We present a significant breakthrough in 3D shape generation by scaling it to unprecedented dimensions. Through the adaptation of the Auto-Regressive model and the utilization of l…

cs.CV2023★ 62 cited

SegGPT: Segmenting Everything In Context

Xinlong Wang, Xiaosong Zhang, Yue Cao +3

We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates di…

cs.CV2023★ 212 cited

EVA-02: A Visual Representation for Neon Genesis

Yuxin Fang, Quan Sun, Xinggang Wang +3

We launch EVA-02, a next-generation Transformer-based visual representation pre-trained to reconstruct strong and robust language-aligned vision features via masked image modeling.…

cs.CV2023★ 1 cited

Hard-aware Instance Adaptive Self-training for Unsupervised Cross-domain Semantic Segmentation

Chuang Zhu, Kebin Liu, Wenqi Tang +3

The divergence between labeled training data and unlabeled testing data is a significant challenge for recent deep learning models. Unsupervised domain adaptation (UDA) attempts to…