315 citations · 539 across the 11 of their papers we have counts for
7 papers · 1 filter
CCQA: A New Web-Scale Question Answering Dataset for Model Pre-Training
Patrick Huber, Armen Aghajanyan, Barlas Oğuz +4
With the rise of large-scale pre-trained language models, open-domain question-answering (ODQA) has become an important research topic in NLP. Based on the popular pre-training fin…
Salient Phrase Aware Dense Retrieval: Can a Dense Retriever Imitate a Sparse One?
Xilun Chen, Kushal Lakhotia, Barlas Oğuz +6
Despite their recent popularity and well-known advantages, dense retrievers still lag behind sparse methods such as BM25 in their ability to reliably match salient phrases and rare…
Domain-matched Pre-training Tasks for Dense Retrieval
Barlas Oğuz, Kushal Lakhotia, Anchit Gupta +8
Pre-training on larger datasets with ever increasing model size is now a proven recipe for increased performance across almost all NLP tasks. A notable exception is information ret…
EASE: Extractive-Abstractive Summarization with Explanations
Haoran Li, Arash Einolghozati, Srinivasan Iyer +4
Current abstractive summarization systems outperform their extractive counterparts, but their widespread adoption is inhibited by the inherent lack of interpretability. To achieve…
El Volumen Louder Por Favor: Code-switching in Task-oriented Semantic Parsing
Arash Einolghozati, Abhinav Arora, Lorena Sainz-Maza Lecanda +2
Being able to parse code-switched (CS) utterances, such as Spanish+English or Hindi+English, is essential to democratize task-oriented semantic parsing systems for certain locales.…
Muppet: Massive Multi-task Representations with Pre-Finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava +3
We propose pre-finetuning, an additional large-scale learning stage between language model pre-training and fine-tuning. Pre-finetuning is massively multi-task learning (around 50…