29 citations · 45 across the 6 of their papers we have counts for
13 papers
What Makes Reading Comprehension Questions Difficult?
Saku Sugawara, Nikita Nangia, Alex Warstadt +1
For a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-…
NOPE: A Corpus of Naturally-Occurring Presuppositions in English
Alicia Parrish, Sebastian Schuster, Alex Warstadt +5
Understanding language requires grasping not only the overtly stated content, but also making inferences about things that were left unsaid. These inferences include presupposition…
What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
Nikita Nangia, Saku Sugawara, Harsh Trivedi +3
Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding…
Does Putting a Linguist in the Loop Improve NLU Data Collection?
Alicia Parrish, William Huang, Omar Agha +7
Many crowdsourced NLP datasets contain systematic gaps and biases that are identified only after data collection is complete. Identifying these issues from early data samples durin…
CLiMP: A Benchmark for Chinese Language Model Evaluation
Beilei Xiang, Changbing Yang, Yu Li +2
Linguistically informed analyses of language models (LMs) contribute to the understanding and improvement of these models. Here, we introduce the corpus of Chinese linguistic minim…
When Do You Need Billions of Words of Pretraining Data?
Yian Zhang, Alex Warstadt, Haau-Sing Li +1
NLP is currently dominated by general-purpose pretrained language models like RoBERTa, which achieve strong performance on NLU tasks through pretraining on billions of words. But w…