most citedMeasuring Progress on Scalable Oversight for Large Language Models

35 citations · 35 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2022

Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions

Alicia Parrish, Harsh Trivedi, Nikita Nangia +4

The use of language-model-based question-answering systems to aid humans in completing difficult tasks is limited, in part, by the unreliability of the text these systems generate.…

cs.CL2022

Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions

Alicia Parrish, Harsh Trivedi, Ethan Perez +4

Current QA systems can generate reasonable-sounding yet false answers without explanation or evidence for the generated answer, which is especially problematic when humans cannot r…

cs.CL2022

What Makes Reading Comprehension Questions Difficult?

Saku Sugawara, Nikita Nangia, Alex Warstadt +1

For a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-…

cs.CL2021

NOPE: A Corpus of Naturally-Occurring Presuppositions in English

Alicia Parrish, Sebastian Schuster, Alex Warstadt +5

Understanding language requires grasping not only the overtly stated content, but also making inferences about things that were left unsaid. These inferences include presupposition…

cs.CL2021

Comparing Test Sets with Item Response Theory

Clara Vania, Phu Mon Htut, William Huang +6

Recent years have seen numerous NLP datasets introduced to evaluate the performance of fine-tuned models on natural language understanding tasks. Recent results from large pretrain…

cs.CL2021

What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?

Nikita Nangia, Saku Sugawara, Harsh Trivedi +3

Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding…