565 citations · 1.4k across the 27 of their papers we have counts for
5 papers · 1 filter
QuALITY: Question Answering with Long Input Texts, Yes!
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi +8
To enable building and testing models on long-document comprehension, we introduce QuALITY, a multiple-choice QA dataset with context passages in English that have an average lengt…
Adversarially Constructed Evaluation Sets Are More Challenging, but May Not Be Fair
Jason Phang, Angelica Chen, William Huang +1
More capable language models increasingly saturate existing task benchmarks, in some cases outperforming humans. This has left little headroom with which to measure further progres…
BBQ: A Hand-Built Bias Benchmark for Question Answering
Alicia Parrish, Angelica Chen, Nikita Nangia +5
It is well documented that NLP models learn social biases, but little work has been done on how these biases manifest in model outputs for applied tasks like question answering (QA…
Fine-Tuned Transformers Show Clusters of Similar Representations Across Layers
Jason Phang, Haokun Liu, Samuel R. Bowman
Despite the success of fine-tuning pretrained language encoders like BERT for downstream natural language understanding (NLU) tasks, it is still poorly understood how neural networ…
Comparing Test Sets with Item Response Theory
Clara Vania, Phu Mon Htut, William Huang +6
Recent years have seen numerous NLP datasets introduced to evaluate the performance of fine-tuned models on natural language understanding tasks. Recent results from large pretrain…