565 citations · 813 across the 12 of their papers we have counts for
Showing 2021Show all
2 papers · 1 filter
cs.CL2021
BBQ: A Hand-Built Bias Benchmark for Question Answering
Alicia Parrish, Angelica Chen, Nikita Nangia +5
It is well documented that NLP models learn social biases, but little work has been done on how these biases manifest in model outputs for applied tasks like question answering (QA…
cs.CL2021
Comparing Test Sets with Item Response Theory
Clara Vania, Phu Mon Htut, William Huang +6
Recent years have seen numerous NLP datasets introduced to evaluate the performance of fine-tuned models on natural language understanding tasks. Recent results from large pretrain…