2 papers
cs.CL2026
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
Longwei Cong, Sonja Hahn, Sebastian Gombert +3
Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa. However, these metrics prov…
cs.CL2026
Confidence Estimation in Automatic Short Answer Grading with LLMs
Longwei Cong, Sonja Hahn, Sebastian Gombert +3
Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fine-tuning, while also enabli…