1 citations · 1 across the 10 of their papers we have counts for
11 papers · 1 filter
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck +57
As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of…
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
Ivo Petrov, Jasper Dekoninck, Dimitar I. Dimitrov +1
Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correctness alone is not sufficient…
Optimizing the Cost-Quality Tradeoff of Agentic Theorem Provers in Lean
Kári Rögnvaldsson, Chenhao Sun, Jasper Dekoninck +1
Large language models (LLMs) are increasingly used in workflows for generating formal proofs in Lean. These workflows often decompose problems into smaller lemmas, sample many proo…
Learning from Saturated Data: Signals Beyond Correctness for LLM Training
Hanno Hiss, Jasper Dekoninck, Martin Vechev
The growing capabilities of large language models (LLMs) have led to the saturation of many benchmarks and training datasets used to improve them. Motivated by this, we investigate…
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Jasper Dekoninck, Nikola JovanoviÄ, Tim Gehrunger +4
Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often na…
The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
Jasper Dekoninck, Ivo Petrov, Kristian Minchev +13
In recent months, large language models (LLMs) have made significant progress in mathematical proof generation, but further advancement is hindered by the lack of a large-scale, hi…