22 citations · 24 across the 4 of their papers we have counts for
1 paper · 1 filter
Zachary Ankner, Naomi Saphra, Davis Blalock +2
Most works on transformers trained with the Masked Language Modeling (MLM) objective use the original BERT model's fixed masking rate of 15%. We propose to instead dynamically sche…