4 papers · 1 filter
Interpretability from the Ground Up: Stakeholder-Centric Design of Automated Scoring in Educational Assessments
Yunsung Kim, Mike Hardy, Joseph Tey +2
AI-driven automated scoring systems offer scalable and efficient means of evaluating complex student-generated responses. Yet, despite increasing demand for transparency and interp…
Autoscoring Anticlimax: A Meta-analytic Understanding of AI's Short-answer Shortcomings and Wording Weaknesses
Michael Hardy
Automated short-answer scoring lags other LLM applications. We meta-analyze 890 culminating results across a systematic review of LLM short-answer scoring studies, modeling the tra…
Measuring Teaching with LLMs
Michael Hardy
Objective and scalable measurement of teaching quality is a persistent challenge in education. While Large Language Models (LLMs) offer potential, general-purpose models have strug…
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
Michael Hardy
"Gold" and "ground truth" human-mediated labels have error. The effects of this error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias,…