activity
20172022
most citedInferring Generative Model Structure with Static Analysis

36 citations · 70 across the 5 of their papers we have counts for

collaborators

7 papers

cs.LG202213 cited

Foundation Transformers

Hongyu Wang, Shuming Ma, Shaohan Huang +12

A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different imp…

cs.LG202219 cited

METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals

Payal Bajaj, Chenyan Xiong, Guolin Ke +7

We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training…

cs.CL20222 cited

Pretraining Text Encoders with Adversarial Mixture of Training Signal Generators

Yu Meng, Chenyan Xiong, Payal Bajaj +4

We present a new framework AMOS that pretrains text encoders with an Adversarial learning curriculum via a Mixture Of Signals from multiple auxiliary generators. Following ELECTRA-…

cs.CL2021

Language Scaling for Universal Suggested Replies Model

Qianlan Ying, Payal Bajaj, Budhaditya Deb +7

We consider the problem of scaling automated suggested replies for Outlook email system to multiple languages. Faced with increased compute requirements and low resources for langu…

cs.CL2021

COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining

Yu Meng, Chenyan Xiong, Payal Bajaj +4

We present a self-supervised learning framework, COCO-LM, that pretrains Language Models by COrrecting and COntrasting corrupted text sequences. Following ELECTRA-style pretraining…

cs.SI2018

Embedding Logical Queries on Knowledge Graphs

William L. Hamilton, Payal Bajaj, Marinka Zitnik +2

Learning low-dimensional embeddings of knowledge graphs is a powerful approach used to predict unobserved or missing edges between entities. However, an open challenge in this area…