195 citations · 221 across the 4 of their papers we have counts for
5 papers
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
Yuxin Fang, Wen Wang, Binhui Xie +6
We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to re…
You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
Yuxin Fang, Bencheng Liao, Xinggang Wang +5
Can Transformer perform 2D object- and region-level recognition from a pure sequence-to-sequence perspective with minimal knowledge about the 2D spatial structure? To answer this q…
Instances as Queries
Yuxin Fang, Shusheng Yang, Xinggang Wang +5
Recently, query based object detection frameworks achieve comparable performance with previous state-of-the-art object detectors. However, how to fully leverage such frameworks to…
Crossover Learning for Fast Online Video Instance Segmentation
Shusheng Yang, Yuxin Fang, Xinggang Wang +5
Modeling temporal visual context across frames is critical for video instance segmentation (VIS) and other video understanding tasks. In this paper, we propose a fast online VIS mo…
Diversity Transfer Network for Few-Shot Learning
Mengting Chen, Yuxin Fang, Xinggang Wang +6
Few-shot learning is a challenging task that aims at training a classifier for unseen classes with only a few training examples. The main difficulty of few-shot learning lies in th…