164 citations · 414 across the 18 of their papers we have counts for
Showing 2021Show all
2 papers · 1 filter
cs.CL2021
Learning to Sample Replacements for ELECTRA Pre-Training
Yaru Hao, Li Dong, Hangbo Bao +2
ELECTRA pretrains a discriminator to detect replaced tokens, where the replacements are sampled from a generator trained with masked language modeling. Despite the compelling perfo…
cs.CL2021
Knowledge Neurons in Pretrained Transformers
Damai Dai, Li Dong, Yaru Hao +3
Large-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus. In this paper, we present preliminary studies on how f…