41 citations · 43 across the 2 of their papers we have counts for
4 papers
Refiner: Refining Self-attention for Vision Transformers
Daquan Zhou, Yujun Shi, Bingyi Kang +6
Vision Transformers (ViTs) have shown competitive accuracy in image classification tasks compared with CNNs. Yet, they generally require much more data for model pre-training. Most…
All Tokens Matter: Token Labeling for Training Better Vision Transformers
Zihang Jiang, Qibin Hou, Li Yuan +5
In this paper, we present token labeling -- a new training objective for training high-performance vision transformers (ViTs). Different from the standard training objective of ViT…
Landau Levels and van der Waals Interfaces of Acoustics in Moiré Phononic Lattices
Shengjie Zheng, Jie Zhang, Guiju Duan +4
Moiré lattices which consist of parallel but staggered periodic lattices have been extensively explored due to their salient physical properties, such as van Hove singularities[1,…
3D Face Reconstruction from A Single Image Assisted by 2D Face Images in the Wild
Xiaoguang Tu, Jian Zhao, Zihang Jiang +6
3D face reconstruction from a single 2D image is a challenging problem with broad applications. Recent methods typically aim to learn a CNN-based 3D face model that regresses coeff…