Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions
Guangxiang Zhao, Qilong Shi, Xusen Xiao +13
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from…
cs.AI2025
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
Lin Sun, Weihong Lin, Jinzhu Wu +8
Reasoning models represented by the Deepseek-R1-Distill series have been widely adopted by the open-source community due to their strong performance in mathematics, science, progra…