2 papers
cs.AI2026
ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions
Guangxiang Zhao, Qilong Shi, Xusen Xiao +13
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from…
cs.AI2026
Thinking with Reasoning Skills: Fewer Tokens, More Accuracy
Guangxiang Zhao, Qilong Shi, Xusen Xiao +3
Reasoning LLMs often spend substantial tokens on long intermediate reasoning traces (e.g., chain-of-thought) when solving new problems. We propose to summarize and store reusable r…