4 citations · 9 across the 13 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
Chuhan Shi, Xiaoquan Ren, Sicheng Song +3
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct t…
cs.AI2026
TeachArena: Are Language Agents Ready for Realistic Teaching Work?
Zixin Chen, Peng Liu, Rui Sheng +6
Language agents are increasingly deployed in professional workflows, yet tutoring remains a high-stakes capability that existing evaluations only partially capture. Effective tutor…