19 citations · 53 across the 8 of their papers we have counts for
16 papers
METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals
Payal Bajaj, Chenyan Xiong, Guolin Ke +7
We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training…
Pretraining Text Encoders with Adversarial Mixture of Training Signal Generators
Yu Meng, Chenyan Xiong, Payal Bajaj +4
We present a new framework AMOS that pretrains text encoders with an Adversarial learning curriculum via a Mixture Of Signals from multiple auxiliary generators. Following ELECTRA-…
Neural Approaches to Conversational Information Retrieval
Jianfeng Gao, Chenyan Xiong, Paul Bennett +1
A conversational information retrieval (CIR) system is an information retrieval (IR) system with a conversational interface which allows users to interact with the system to seek i…
Zero-Shot Dense Retrieval with Momentum Adversarial Domain Invariant Representations
Ji Xin, Chenyan Xiong, Ashwin Srinivasan +3
Dense retrieval (DR) methods conduct text retrieval by first encoding texts in the embedding space and then matching them by nearest neighbor search. This requires strong locality…
Keep it Simple: Unsupervised Simplification of Multi-Paragraph Text
Philippe Laban, Tobias Schnabel, Paul Bennett +1
This work presents Keep it Simple (KiS), a new approach to unsupervised text simplification which learns to balance a reward across three properties: fluency, salience and simplici…
Domain-Specific Pretraining for Vertical Search: Case Study on Biomedical Literature
Yu Wang, Jinchao Li, Tristan Naumann +12
Information overload is a prevalent challenge in many high-value domains. A prominent case in point is the explosion of the biomedical literature on COVID-19, which swelled to hund…