2 papers
cs.CL2026
General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
Junlin Liu, Shengnan An, Shuang Zhou +10
Contemporary large language models (LLMs) have demonstrated remarkable reasoning capabilities, particularly in specialized domains like mathematics and physics. However, their abil…
cs.CL2025
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
Shengnan An, Xunliang Cai, Xuezhi Cao +8
We present AMO-Bench, an Advanced Mathematical reasoning benchmark with Olympiad level or even higher difficulty, comprising 50 human-crafted problems. Existing benchmarks have wid…