activity
20192022
most citedYou Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection

195 citations · 221 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV202223 cited

EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

Yuxin Fang, Wen Wang, Binhui Xie +6

We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to re…

cs.CV2021195 cited

You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection

Yuxin Fang, Bencheng Liao, Xinggang Wang +5

Can Transformer perform 2D object- and region-level recognition from a pure sequence-to-sequence perspective with minimal knowledge about the 2D spatial structure? To answer this q…

cs.CV2021

Instances as Queries

Yuxin Fang, Shusheng Yang, Xinggang Wang +5

Recently, query based object detection frameworks achieve comparable performance with previous state-of-the-art object detectors. However, how to fully leverage such frameworks to…

cs.CV20213 cited

Crossover Learning for Fast Online Video Instance Segmentation

Shusheng Yang, Yuxin Fang, Xinggang Wang +5

Modeling temporal visual context across frames is critical for video instance segmentation (VIS) and other video understanding tasks. In this paper, we propose a fast online VIS mo…

cs.CV2019

Diversity Transfer Network for Few-Shot Learning

Mengting Chen, Yuxin Fang, Xinggang Wang +6

Few-shot learning is a challenging task that aims at training a classifier for unseen classes with only a few training examples. The main difficulty of few-shot learning lies in th…