30 citations · 137 across the 12 of their papers we have counts for
18 papers
Beyond Masking: Demystifying Token-Based Pre-Training for Vision Transformers
Yunjie Tian, Lingxi Xie, Jiemin Fang +6
The past year has witnessed a rapid development of masked image modeling (MIM). MIM is mostly built upon the vision transformers, which suggests that self-supervised visual represe…
P2P-Loc: Point to Point Tiny Person Localization
Xuehui Yu, Di Wu, Qixiang Ye +2
Bounding-box annotation form has been the most frequently used method for visual object localization tasks. However, bounding-box annotation relies on a large amount of precisely a…
Long-tailed Distribution Adaptation
Zhiliang Peng, Wei Huang, Zonghao Guo +3
Recognizing images with long-tailed distributions remains a challenging problem while there lacks an interpretable mechanism to solve this problem. In this study, we formulate Long…
GraFormer: Graph Convolution Transformer for 3D Pose Estimation
Weixi Zhao, Yunjie Tian, Qixiang Ye +2
Exploiting relations among 2D joints plays a crucial role yet remains semi-developed in 2D-to-3D pose estimation. To alleviate this issue, we propose GraFormer, a novel transformer…
Anti-aliasing Semantic Reconstruction for Few-Shot Semantic Segmentation
Binghao Liu, Yao Ding, Jianbin Jiao +2
Encouraging progress in few-shot semantic segmentation has been made by leveraging features learned upon base classes with sufficient training data to represent novel classes with…
Conformer: Local Features Coupling Global Representations for Visual Recognition
Zhiliang Peng, Wei Huang, Shanzhi Gu +4
Within Convolutional Neural Network (CNN), the convolution operations are good at extracting local features but experience difficulty to capture global representations. Within visu…