papers

Publications (9)

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

cs.CL2026

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger +4

Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often na…

math.AG2026

The Classification of the Stable Marked Reduction of Genus 2 Curves in Residue Characteristic 2

Tim Gehrunger

Consider a hyperelliptic curve of genus over a field of characteristic zero. After extending we can view it as a marked curve with its Weierstrass points. We classi…

math.AG2025

Computing the Stable Reduction of Hyperelliptic Curves in Residue Characteristic 2

Tim Gehrunger

Consider a hyperelliptic curve of genus over a field of characteristic zero. After extending we can view it as a marked curve with its Weierstrass points. We pro…

math.HO2026

Benchmarks in Leipzig

Andrei Balakin, Miklós Bóna, Marie-Charlotte Brandenburg +45

Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3…

cs.AI2026

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck +4

Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tai…