Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
Daniel Kaiser, Arnoldo Frigessi, Ali Ramezani-Kebrya +1
As reasoning LLMs increasingly trade tokens for accuracy through deliberation, search, and self-correction, a single accuracy score can no longer tell whether those tokens buy usef…
cs.CL2025
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
Daniel Kaiser, Arnoldo Frigessi, Ali Ramezani-Kebrya +1
Current benchmarks for long-context reasoning in Large Language Models (LLMs) often blur critical factors like intrinsic task complexity, distractor interference, and task length.…