Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Measuring Teaching with LLMs
Michael Hardy
Objective and scalable measurement of teaching quality is a persistent challenge in education. While Large Language Models (LLMs) offer potential, general-purpose models have strug…
cs.CL2024
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
Michael Hardy
"Gold" and "ground truth" human-mediated labels have error. The effects of this error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias,…