Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Tianyou Wang, Chongyang Gao, Kezhen Chen +7
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the…
cs.AI2026
Classroom Final Exam: An Instructor-Tested Reasoning Benchmark
Chongyang Gao, Diji Yang, Shuyan Zhou +4
We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench…