Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SciEval: A Benchmark for Automatic Evaluation of K-12 Science Instructional Materials
Zhaohui Li, Peng He, Zhiyuan Chen +4
The need to evaluate instructional materials for K-12 science education has become increasingly important, as more educators use generative AI to create instructional materials. Ho…
cs.AI2026
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
Hang Li, Kaiqi Yang, Xianxuan Long +9
The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptabili…