3 papers
cs.AI2026
Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics
Moritz Firsching, Paul Lezeau, Salvatore Mercuri +8
As automated reasoning systems advance rapidly, there is a growing need for research-level formal mathematical problems to accurately evaluate their capabilities. To address this,…
math.CO2025
A Generalisation of Sperner's Theorem Using Weighted Chains
Yaël Dillies, Matthew Johnson, Aleksandra Kowalska
We find the (unique) largest subset of such that it contains no two elements, one of which is coordinatewise greater than the other, but strictly greater on at most…
cs.AI2025
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
Junqi Liu, Xiaohan Lin, Jonas Bayer +12
Neurosymbolic approaches integrating large language models with formal reasoning have recently achieved human-level performance on mathematics competition problems in algebra, geom…