7 papers
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck +4
Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tai…
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck +57
As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of…
Benchmarks in Leipzig
Andrei Balakin, Miklós Bóna, Marie-Charlotte Brandenburg +45
Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3…
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Jasper Dekoninck, Nikola JovanoviÄ, Tim Gehrunger +4
Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often na…
The Classification of the Stable Marked Reduction of Genus 2 Curves in Residue Characteristic 2
Tim Gehrunger
Consider a hyperelliptic curve of genus over a field of characteristic zero. After extending we can view it as a marked curve with its Weierstrass points. We classi…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…