3 papers
cs.AI2026
SciEval: A Benchmark for Automatic Evaluation of K-12 Science Instructional Materials
Zhaohui Li, Peng He, Zhiyuan Chen +4
The need to evaluate instructional materials for K-12 science education has become increasingly important, as more educators use generative AI to create instructional materials. Ho…
cs.CY2026
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
Peng He, Zhaohui Li, Zeyuan Wang +2
Designing high-quality, standards-aligned instructional materials for K--12 science is time-consuming and expertise-intensive. This study examines what human experts notice when re…
cs.AI2025
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
Chenhui Xu, Dancheng Liu, Jiajie Li +3
Recent advancements in cognitive science and multi-round reasoning techniques for Large Language Models (LLMs) suggest that iterative thinking processes improve problem-solving per…