1 paper
Qihuang Zhong, Liang Ding, Juhua Liu +4
Token dropping is a recently-proposed strategy to speed up the pretraining of masked language models, such as BERT, by skipping the computation of a subset of the input tokens at s…