activity
20192022
most citedAligned Cross Entropy for Non-Autoregressive Machine Translation

68 citations · 113 across the 4 of their papers we have counts for

collaborators

9 papers

cs.CL202242 cited

CM3: A Causal Masked Multimodal Model of the Internet

Armen Aghajanyan, Bernie Huang, Candace Ross +8

We introduce CM3, a family of causally masked generative models trained over a large corpus of structured multi-modal documents that can contain both text and image tokens. Our new…

cs.CL20211 cited

Domain-matched Pre-training Tasks for Dense Retrieval

Barlas Oğuz, Kushal Lakhotia, Anchit Gupta +8

Pre-training on larger datasets with ever increasing model size is now a proven recipe for increased performance across almost all NLP tasks. A notable exception is information ret…

cs.CL2021

NeurIPS 2020 EfficientQA Competition: Systems, Analyses and Lessons Learned

Sewon Min, Jordan Boyd-Graber, Chris Alberti +50

We review the EfficientQA competition from NeurIPS 2020. The competition focused on open-domain question answering (QA), where systems take natural language questions as input and…

cs.CL2021

Multi-task Retrieval for Knowledge-Intensive Tasks

Jean Maillard, Vladimir Karpukhin, Fabio Petroni +4

Retrieving relevant contexts from a large corpus is a crucial step for tasks such as open-domain question answering and fact checking. Although neural retrieval outperforms traditi…

cs.CL2020

Joint Verification and Reranking for Open Fact Checking Over Tables

Michael Schlichtkrull, Vladimir Karpukhin, Barlas Oğuz +3

Structured information is an important knowledge source for automatic verification of factual claims. Nevertheless, the majority of existing research into this task has focused on…

cs.CL2020

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Patrick Lewis, Ethan Perez, Aleksandra Piktus +9

Large pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. Howe…