5 papers
Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions
Alicia Parrish, Harsh Trivedi, Nikita Nangia +4
The use of language-model-based question-answering systems to aid humans in completing difficult tasks is limited, in part, by the unreliability of the text these systems generate.…
Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions
Alicia Parrish, Harsh Trivedi, Ethan Perez +4
Current QA systems can generate reasonable-sounding yet false answers without explanation or evidence for the generated answer, which is especially problematic when humans cannot r…
NOPE: A Corpus of Naturally-Occurring Presuppositions in English
Alicia Parrish, Sebastian Schuster, Alex Warstadt +5
Understanding language requires grasping not only the overtly stated content, but also making inferences about things that were left unsaid. These inferences include presupposition…
Does Putting a Linguist in the Loop Improve NLU Data Collection?
Alicia Parrish, William Huang, Omar Agha +7
Many crowdsourced NLP datasets contain systematic gaps and biases that are identified only after data collection is complete. Identifying these issues from early data samples durin…
Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs
Alex Warstadt, Yu Cao, Ioana Grosu +13
Though state-of-the-art sentence representation models can perform tasks requiring significant knowledge of grammar, it is an open question how best to evaluate their grammatical k…