activity
20182025
most citedDynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

311 citations · 608 across the 18 of their papers we have counts for

collaborators
Showing 2021Show all

5 papers · 1 filter

cs.CV2021★ 38 cited

DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting

Yongming Rao, Wenliang Zhao, Guangyi Chen +5

Recent progress has shown that large-scale pre-training using contrastive image-text pairs can be a promising alternative for high-quality visual representation learning from natur…

cs.CV2021★ 1 cited

Group-aware Contrastive Regression for Action Quality Assessment

Xumin Yu, Yongming Rao, Wenliang Zhao +2

Assessing action quality is challenging due to the subtle differences between videos and large variations in scores. Most existing approaches tackle this problem by regressing a qu…

cs.CV2021★ 1 cited

Towards Interpretable Deep Metric Learning with Structural Matching

Wenliang Zhao, Yongming Rao, Ziyi Wang +2

How do the neural networks distinguish two images? It is of critical importance to understand the matching mechanism of deep models for developing reliable intelligent systems for…

cs.CV2021★ 19 cited

Global Filter Networks for Image Classification

Yongming Rao, Wenliang Zhao, Zheng Zhu +2

Recent advances in self-attention and pure multi-layer perceptrons (MLP) models for vision have shown great potential in achieving promising performance with fewer inductive biases…

cs.CV2021★ 311 cited

DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

Yongming Rao, Wenliang Zhao, Benlin Liu +3

Attention is sparse in vision transformers. We observe the final prediction in vision transformers is only based on a subset of most informative tokens, which is sufficient for acc…