45 citations · 79 across the 3 of their papers we have counts for
1 paper · 1 filter
Yoav Levine, Barak Lenz, Opher Lieber +4
Masking tokens uniformly at random constitutes a common flaw in the pretraining of Masked Language Models (MLMs) such as BERT. We show that such uniform masking allows an MLM to mi…