18 citations · 18 across the 2 of their papers we have counts for
3 papers
cs.AI2026
Human vs Machine Mathematical Difficulty on Project Euler: An Experimental Analysis
David Holmes, Johannes Schmitt
We study how the effort and success probability of frontier AI systems scale with human difficulty on problems from Project Euler, an online platform of computational mathematics p…
cs.LG2026★ 18 cited
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
cs.CC2025
Parameterised Holant Problems
Panagiotis Aivasiliotis, Andreas Göbel, Marc Roth +1
We investigate the complexity of parameterised holant problems p- for families of signatures . The parameterised holant framework was int…