1 paper
Antoine Peyronnet, Fabian Gloeckle, Amaury Hayat
We present a new approach for benchmarking Large Language Model (LLM) capabilities on research-level mathematics. Existing benchmarks largely rely on static, hand-curated sets of c…