4 papers
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Chenguang Wang, Ming Li, Xinyue Zeng +4
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on c…
Comparison of Scoring Rationales Between Large Language Models and Human Raters
Haowei Hua, Hong Jiao, Dan Song
Advances in automated scoring are closely aligned with advances in machine-learning and natural-language-processing techniques. With recent progress in large language models (LLMs)…
Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory
Dan Song, Won-Chan Lee, Hong Jiao
This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Culture Exam. Using generalizabilit…
Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model
Hong Jiao, Dan Song, Won-Chan Lee
Large language models (LLMs) have been widely explored for automated scoring in low-stakes assessment to facilitate learning and instruction. Empirical evidence related to which LL…