35 citations · 35 across the 5 of their papers we have counts for
6 papers · 1 filter
Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions
Alicia Parrish, Harsh Trivedi, Nikita Nangia +4
The use of language-model-based question-answering systems to aid humans in completing difficult tasks is limited, in part, by the unreliability of the text these systems generate.…
Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions
Alicia Parrish, Harsh Trivedi, Ethan Perez +4
Current QA systems can generate reasonable-sounding yet false answers without explanation or evidence for the generated answer, which is especially problematic when humans cannot r…
What Makes Reading Comprehension Questions Difficult?
Saku Sugawara, Nikita Nangia, Alex Warstadt +1
For a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-…
NOPE: A Corpus of Naturally-Occurring Presuppositions in English
Alicia Parrish, Sebastian Schuster, Alex Warstadt +5
Understanding language requires grasping not only the overtly stated content, but also making inferences about things that were left unsaid. These inferences include presupposition…
Comparing Test Sets with Item Response Theory
Clara Vania, Phu Mon Htut, William Huang +6
Recent years have seen numerous NLP datasets introduced to evaluate the performance of fine-tuned models on natural language understanding tasks. Recent results from large pretrain…
What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
Nikita Nangia, Saku Sugawara, Harsh Trivedi +3
Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding…