1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.AI2026
TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models
Zheng Li, Siyao Song, Jingyuan Ma +4
Large language models (LLMs) show promise as teaching assistants, yet their teaching capability remains insufficiently evaluated. Existing benchmarks mainly focus on problem-solvin…
cs.CL2025★ 1 cited
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Minghao Li, Ying Zeng, Zhihao Cheng +2
The advent of Deep Research agents has substantially reduced the time required for conducting extensive research tasks. However, these tasks inherently demand rigorous standards of…