activity
20172023
most citedDeepViT: Towards Deeper Vision Transformer

349 citations · 666 across the 15 of their papers we have counts for

collaborators

16 papers

cs.CV20212 cited

Video Salient Object Detection via Contrastive Features and Attention Modules

Yi-Wen Chen, Xiaojie Jin, Xiaohui Shen +1

Video salient object detection aims to find the most visually distinctive objects in a video. To explore the temporal dependencies, existing methods usually resort to recurrent neu…

cs.CV20213 cited

HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers

Mingyu Ding, Xiaochen Lian, Linjie Yang +4

High-resolution representations (HR) are essential for dense prediction tasks such as segmentation, detection, and pose estimation. Learning HR representations is typically ignored…

cs.CV202141 cited

Refiner: Refining Self-attention for Vision Transformers

Daquan Zhou, Yujun Shi, Bingyi Kang +6

Vision Transformers (ViTs) have shown competitive accuracy in image classification tasks compared with CNNs. Yet, they generally require much more data for model pre-training. Most…

cs.LG2021

One Backward from Ten Forward, Subsampling for Large-Scale Deep Learning

Chaosheng Dong, Xiaojie Jin, Weihao Gao +5

Deep learning models in large-scale machine learning systems are often continuously trained with enormous data from production environments. The sheer volume of streaming training…

cs.CV2021349 cited

DeepViT: Towards Deeper Vision Transformer

Daquan Zhou, Bingyi Kang, Xiaojie Jin +5

Vision transformers (ViTs) have been successfully applied in image classification tasks recently. In this paper, we show that, unlike convolution neural networks (CNNs)that can be…

cs.CV2021

All Tokens Matter: Token Labeling for Training Better Vision Transformers

Zihang Jiang, Qibin Hou, Li Yuan +5

In this paper, we present token labeling -- a new training objective for training high-performance vision transformers (ViTs). Different from the standard training objective of ViT…