3 papers
math.HO2026
Benchmarks in Leipzig
Andrei Balakin, Miklós Bóna, Marie-Charlotte Brandenburg +45
Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3…
cs.CL2026
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
Guijin Son, Seungone Kim, Catherine Arnett +73
Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM…
math.NT2023
Testing local-global divisibility at a stable set
Alexander B. Ivanov, Laura Paladino
We show that the local-global divisibility in commutative algebraic groups defined over number fields can be tested on sets of primes of arbitrary small density, i.e. stable and pe…