1 paper
Yu Lin, Zhecheng An, Peihao Wu +1
Though achieving impressive results on many NLP tasks, the BERT-like masked language models (MLM) encounter the discrepancy between pre-training and inference. In light of this gap…