Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize
Jiaxin Yuan, Connor Martinez Lockhart, Xiaoyu Liu +11
Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow ra…
cs.CL2026
Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models
Renuka Oladri, Niveda Jawahar, Abdirisak Mohamed
Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a token budget (converged) or exhau…