2 papers
cs.PL2025
L0-Reasoning Bench: Evaluating Procedural Correctness in Language Models via Simple Program Execution
Simeng Sun, Cheng-Ping Hsieh, Faisal Ladhak +3
Complex reasoning tasks often rely on the ability to consistently and accurately apply simple rules across incremental steps, a foundational capability which we term "level-0" reas…
cs.AI2025
Reasoning Inconsistencies and How to Mitigate Them in Deep Learning
Erik Arakelyan
The recent advancements in Deep Learning models and techniques have led to significant strides in performance across diverse tasks and modalities. However, while the overall capabi…