3 papers
cs.LG2026
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye +48
As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that st…
cs.AI2026
Inference-Time Diversity in RL-Trained Lean Theorem Provers: A Diagnostic Study
Zachary Burton
RL-trained Lean theorem provers mode-collapse at inference time: on miniF2F-test with DeepSeek-Prover-V1.5-RL, doubling the i.i.d.\ sampling budget from to produc…
math.PR2025
Poissonization-based collision threshold derivation for random walks on lattices
Zachary Burton
In this expository note, we give a short derivation of the expected number of collisions between two independent simple random walkers on integer lattices. Adapting a Poissonizatio…