4 citations · 4 across the 1 of their papers we have counts for
1 paper
Andy Wagner, Tiyasa Mitra, Mrinal Iyer +2
Masked language modeling (MLM) pre-training models such as BERT corrupt the input by replacing some tokens with [MASK] and then train a model to reconstruct the original tokens. Th…