papers
Publications (2)
cs.LG2026
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
cs.CC2022
Refuting Tianrong Lin's arXiv:2110.05942 "Resolution of The Linear-Bounded Automata Question"
Thomas Preu
In the preprint mentioned in the title Mr. Tianrong claims to prove , resolving a longstanding open problem in automata theory called the…