3 papers
cs.AI2026
Human agency in initial human-AI proof formalization workflows
Katherine M. Collins, Simon Frieder, Jonas Bayer +14
For centuries, human mathematicians have written proofs to substantiate their mathematical arguments; yet, the ability to automatically verify the validity of proofs has long been…
cs.LG2026
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
cs.LG2025
Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning
Simon Frieder, Jonas Bayer, Sam Looi +13
The datasets and benchmarks commonly used to train and evaluate the mathematical capabilities of AI-based mathematical copilots (primarily large language models) exhibit several sh…