1 paper
Tao Liu, Ye Lu, Ruohua Zhang +4
Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain-general correctness or depe…