3 papers
cs.AI2025
Fantastic Bugs and Where to Find Them in AI Benchmarks
Sang Truong, Yuheng Tu, Michael Hardy +8
Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of…
cs.CL2025
Measuring Teaching with LLMs
Michael Hardy
Objective and scalable measurement of teaching quality is a persistent challenge in education. While Large Language Models (LLMs) offer potential, general-purpose models have strug…
cs.CL2024
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
Michael Hardy
"Gold" and "ground truth" human-mediated labels have error. The effects of this error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias,…