activity
20162024
most citedHigh-Resolution Representations for Labeling Pixels and Regions

665 citations · 1.1k across the 31 of their papers we have counts for

collaborators

49 papers

cs.CV20242 cited

EVA-X: A Foundation Model for General Chest X-ray Analysis with Self-supervised Learning

Jingfeng Yao, Xinggang Wang, Yuehao Song +5

The diagnosis and treatment of chest diseases play a crucial role in maintaining human health. X-ray examination has become the most common clinical examination means due to its ef…

cs.CV20224 cited

Perceive, Interact, Predict: Learning Dynamic and Static Clues for End-to-End Motion Prediction

Bo Jiang, Shaoyu Chen, Xinggang Wang +7

Motion prediction is highly relevant to the perception of dynamic objects and static map elements in the scenarios of autonomous driving. In this work, we propose PIP, the first en…

cs.CV202223 cited

EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

Yuxin Fang, Wen Wang, Binhui Xie +6

We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to re…

cs.CV20224 cited

Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object Detection

Yuxin Fang, Shusheng Yang, Shijie Wang +3

We present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our t…

cs.CV20222 cited

Temporally Efficient Vision Transformer for Video Instance Segmentation

Shusheng Yang, Xinggang Wang, Yu Li +5

Recently vision transformer has achieved tremendous success on image-level visual recognition tasks. To effectively and efficiently model the crucial temporal information within a…

cs.CV202216 cited

TopFormer: Token Pyramid Transformer for Mobile Semantic Segmentation

Wenqiang Zhang, Zilong Huang, Guozhong Luo +5

Although vision transformers (ViTs) have achieved great success in computer vision, the heavy computational cost hampers their applications to dense prediction tasks such as semant…