299 citations · 322 across the 4 of their papers we have counts for
10 papers
METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals
Payal Bajaj, Chenyan Xiong, Guolin Ke +7
We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training…
Pretraining Text Encoders with Adversarial Mixture of Training Signal Generators
Yu Meng, Chenyan Xiong, Payal Bajaj +4
We present a new framework AMOS that pretrains text encoders with an Adversarial learning curriculum via a Mixture Of Signals from multiple auxiliary generators. Following ELECTRA-…
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Shaden Smith, Mostofa Patwary, Brandon Norick +17
Pretrained general-purpose language models can achieve state-of-the-art accuracies in various natural language processing domains by adapting to downstream tasks via zero-shot, few…
COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining
Yu Meng, Chenyan Xiong, Payal Bajaj +4
We present a self-supervised learning framework, COCO-LM, that pretrains Language Models by COrrecting and COntrasting corrupted text sequences. Following ELECTRA-style pretraining…
Knowledge-Aware Language Model Pretraining
Corby Rosset, Chenyan Xiong, Minh Phan +3
How much knowledge do pretrained language models hold? Recent research observed that pretrained transformers are adept at modeling semantics but it is unclear to what degree they g…
Generic Intent Representation in Web Search
Hongfei Zhang, Xia Song, Chenyan Xiong +4
This paper presents GEneric iNtent Encoder (GEN Encoder) which learns a distributed representation space for user intent in search. Leveraging large scale user clicks from Bing sea…