4 papers
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
Longwei Cong, Sonja Hahn, Sebastian Gombert +3
Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa. However, these metrics prov…
Confidence Estimation in Automatic Short Answer Grading with LLMs
Longwei Cong, Sonja Hahn, Sebastian Gombert +3
Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fine-tuning, while also enabli…
Report on the Scoping Workshop on AI in Science Education Research 2025
Marcus Kubsch, Marit Kastaun, Peter Wulff +21
This report summarizes the outcomes of a two-day international scoping workshop on the role of artificial intelligence (AI) in science education research. As AI rapidly reshapes sc…
Concept Map Assessment Through Structure Classification
LaÃs P. V. Vossen, Isabela Gasparini, Elaine H. T. Oliveira +6
Due to their versatility, concept maps are used in various educational settings and serve as tools that enable educators to comprehend students' knowledge construction. An essentia…