Publications (9)
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Jasper Dekoninck, Nikola JovanoviÄ, Tim Gehrunger +4
Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often na…
The Classification of the Stable Marked Reduction of Genus 2 Curves in Residue Characteristic 2
Tim Gehrunger
Consider a hyperelliptic curve of genus over a field of characteristic zero. After extending we can view it as a marked curve with its Weierstrass points. We classi…
Computing the Stable Reduction of Hyperelliptic Curves in Residue Characteristic 2
Tim Gehrunger
Consider a hyperelliptic curve of genus over a field of characteristic zero. After extending we can view it as a marked curve with its Weierstrass points. We pro…
Benchmarks in Leipzig
Andrei Balakin, Miklós Bóna, Marie-Charlotte Brandenburg +45
Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3…
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck +4
Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tai…