3 papers
cs.CL2021
Learning to Sample Replacements for ELECTRA Pre-Training
Yaru Hao, Li Dong, Hangbo Bao +2
ELECTRA pretrains a discriminator to detect replaced tokens, where the replacements are sampled from a generator trained with masked language modeling. Despite the compelling perfo…
cs.CL2020
Self-Attention Attribution: Interpreting Information Interactions Inside Transformer
Yaru Hao, Li Dong, Furu Wei +1
The great success of Transformer-based models benefits from the powerful multi-head self-attention mechanism, which learns token dependencies and encodes contextual information fro…
cs.CL2019
Visualizing and Understanding the Effectiveness of BERT
Yaru Hao, Li Dong, Furu Wei +1
Language model pre-training, such as BERT, has achieved remarkable results in many NLP tasks. However, it is unclear why the pre-training-then-fine-tuning paradigm can improve perf…