activity
20182022
most citedUniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training

225 citations · 250 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV202214 cited

A Unified View of Masked Image Modeling

Zhiliang Peng, Li Dong, Hangbo Bao +2

Masked image modeling has demonstrated great potential to eliminate the label-hungry problem of training large-scale vision Transformers, achieving impressive performance on variou…

cs.CL202111 cited

s2s-ft: Fine-Tuning Pretrained Transformer Encoders for Sequence-to-Sequence Learning

Hangbo Bao, Li Dong, Wenhui Wang +2

Pretrained bidirectional Transformers, such as BERT, have achieved significant improvements in a wide variety of language understanding tasks, while it is not straightforward to di…

cs.CL2021

Learning to Sample Replacements for ELECTRA Pre-Training

Yaru Hao, Li Dong, Hangbo Bao +2

ELECTRA pretrains a discriminator to detect replaced tokens, where the replacements are sampled from a generator trained with masked language modeling. Despite the compelling perfo…

cs.CL2020

MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers

Wenhui Wang, Hangbo Bao, Shaohan Huang +2

We generalize deep self-attention distillation in MiniLM (Wang et al., 2020) by only using self-attention relation distillation for task-agnostic compression of pretrained Transfor…

cs.CL2020225 cited

UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training

Hangbo Bao, Li Dong, Furu Wei +8

We propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseud…

cs.CL2020

MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers

Wenhui Wang, Furu Wei, Li Dong +3

Pre-trained language models (e.g., BERT (Devlin et al., 2018) and its variants) have achieved remarkable success in varieties of NLP tasks. However, these models usually consist of…