Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AIMO Interpretability Challenge
Michal Štefánik, Philipp Mondorf, Andreas Waldis +11
We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' interna…
cs.AI2025
Prover Agent: An Agent-Based Framework for Formal Mathematical Proofs
Kaito Baba, Chaoran Liu, Shuhei Kurita +1
We present Prover Agent, a novel AI agent for automated theorem proving that integrates large language models (LLMs) with a formal proof assistant, Lean. Prover Agent coordinates a…