activity
20242026
collaborators

5 papers

cs.CL2026

EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading

Zhihao Wu, Linhai Zhang, Taiyi Wang +4

Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student answer. Existing credit-assig…

cs.CL2025

Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time

Jiazheng Li, Yuxiang Zhou, Junru Lu +4

Although preference optimization methods have improved reasoning performance in Large Language Models (LLMs), they often lack transparency regarding why one reasoning outcome is pr…

cs.CL2025

AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment

Jiazheng Li, Artem Bobrov, Runcong Zhao +2

Explainability in automated student answer scoring systems is critical for building trust and enhancing usability among educators. Yet, generating high-quality assessment rationale…

cs.CL2024

An Automated Explainable Educational Assessment System Built on LLMs

Jiazheng Li, Artem Bobrov, David West +2

In this demo, we present AERA Chat, an automated and explainable educational assessment system designed for interactive and visual evaluations of student responses. This system lev…

cs.CL2024

Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring

Jiazheng Li, Hainiu Xu, Zhaoyue Sun +4

Generating rationales that justify scoring decisions has been a promising way to facilitate explainability in automated scoring systems. However, existing methods do not match the…