2 papers
cs.CL2021
Comparing Test Sets with Item Response Theory
Clara Vania, Phu Mon Htut, William Huang +6
Recent years have seen numerous NLP datasets introduced to evaluate the performance of fine-tuned models on natural language understanding tasks. Recent results from large pretrain…
cs.CL2020
Precise Task Formalization Matters in Winograd Schema Evaluations
Haokun Liu, William Huang, Dhara A. Mungra +1
Performance on the Winograd Schema Challenge (WSC), a respected English commonsense reasoning benchmark, recently rocketed from chance accuracy to 89% on the SuperGLUE leaderboard,…