7 citations · 7 across the 2 of their papers we have counts for
2 papers
cs.CL2022
Token Dropping for Efficient BERT Pretraining
Le Hou, Richard Yuanzhe Pang, Tianyi Zhou +4
Transformer-based models generally allocate the same amount of computation for each token in a given sequence. We develop a simple but effective "token dropping" method to accelera…
cs.LG2022★ 7 cited
Auto-scaling Vision Transformers without Training
Wuyang Chen, Wei Huang, Xianzhi Du +3
This work targets automated designing and scaling of Vision Transformers (ViTs). The motivation comes from two pain spots: 1) the lack of efficient and principled methods for desig…