5 papers
EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading
Zhihao Wu, Linhai Zhang, Taiyi Wang +4
Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student answer. Existing credit-assig…
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time
Jiazheng Li, Yuxiang Zhou, Junru Lu +4
Although preference optimization methods have improved reasoning performance in Large Language Models (LLMs), they often lack transparency regarding why one reasoning outcome is pr…
AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment
Jiazheng Li, Artem Bobrov, Runcong Zhao +2
Explainability in automated student answer scoring systems is critical for building trust and enhancing usability among educators. Yet, generating high-quality assessment rationale…
An Automated Explainable Educational Assessment System Built on LLMs
Jiazheng Li, Artem Bobrov, David West +2
In this demo, we present AERA Chat, an automated and explainable educational assessment system designed for interactive and visual evaluations of student responses. This system lev…
Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring
Jiazheng Li, Hainiu Xu, Zhaoyue Sun +4
Generating rationales that justify scoring decisions has been a promising way to facilitate explainability in automated scoring systems. However, existing methods do not match the…