3 papers
cs.LG2026
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye +48
As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that st…
math.CO2025
The structure of -free tournaments
Seokbeom Kim, Taite LaGrange, Mathieu Rundström +2
We extend the list of tournaments for which the complete structural description for tournaments excluding as a subtournament is known. Specifically, let be a…
math.CO2025
An improved upper bound for the multicolour Ramsey number of odd cycles
Maria Axenovich, Wouter Cames van Batenburg, Oliver Janzer +2
We show that the -colour Ramsey number of an odd cycle of length is at most . This proves a conjecture of Fox and is the first improvem…