17 papers
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck +4
Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tai…
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck +57
As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of…
Adaptive Generation of Bias-Eliciting Questions for LLMs
Robin Staab, Jasper Dekoninck, Maximilian Baader +1
Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide. Despite their widespread adoption, growing relia…
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness
Ivo Petrov, Jasper Dekoninck, Dimitar I. Dimitrov +1
Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correctness alone is not sufficient…
Optimizing the Cost-Quality Tradeoff of Agentic Theorem Provers in Lean
Kári Rögnvaldsson, Chenhao Sun, Jasper Dekoninck +1
Large language models (LLMs) are increasingly used in workflows for generating formal proofs in Lean. These workflows often decompose problems into smaller lemmas, sample many proo…
Learning from Saturated Data: Signals Beyond Correctness for LLM Training
Hanno Hiss, Jasper Dekoninck, Martin Vechev
The growing capabilities of large language models (LLMs) have led to the saturation of many benchmarks and training datasets used to improve them. Motivated by this, we investigate…