2 papers
cs.CL2025
Frontier LLMs Still Struggle with Simple Reasoning Tasks
Alan Malek, Jiawei Ge, Nevena Lazic +3
While state-of-the-art large language models (LLMs) demonstrate advanced reasoning capabilities-achieving remarkable performance on challenging competitive math and coding benchmar…
cs.LG2025
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
Kaixuan Huang, Jiacheng Guo, Zihao Li +15
Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieve…