2 citations · 2 across the 1 of their papers we have counts for
1 paper
Micah Carroll, Orr Paradise, Jessy Lin +8
Randomly masking and predicting word tokens has been a successful approach in pre-training language models for a variety of downstream tasks. In this work, we observe that the same…