311 citations · 608 across the 18 of their papers we have counts for
5 papers · 1 filter
DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen +5
Recent progress has shown that large-scale pre-training using contrastive image-text pairs can be a promising alternative for high-quality visual representation learning from natur…
Group-aware Contrastive Regression for Action Quality Assessment
Xumin Yu, Yongming Rao, Wenliang Zhao +2
Assessing action quality is challenging due to the subtle differences between videos and large variations in scores. Most existing approaches tackle this problem by regressing a qu…
Towards Interpretable Deep Metric Learning with Structural Matching
Wenliang Zhao, Yongming Rao, Ziyi Wang +2
How do the neural networks distinguish two images? It is of critical importance to understand the matching mechanism of deep models for developing reliable intelligent systems for…
Global Filter Networks for Image Classification
Yongming Rao, Wenliang Zhao, Zheng Zhu +2
Recent advances in self-attention and pure multi-layer perceptrons (MLP) models for vision have shown great potential in achieving promising performance with fewer inductive biases…
DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu +3
Attention is sparse in vision transformers. We observe the final prediction in vision transformers is only based on a subset of most informative tokens, which is sufficient for acc…