Publications (4)
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Lower bound for deterministic semantic-incremental branching programs solving GEN
Dustin Wehr
We answer a problem posed in (Gál, Koucký, McKenzie 2008) regarding a restricted model of small-space computation, tailored for solving the GEN problem. They define two variants…
Pebbles and Branching Programs for Tree Evaluation
Stephen Cook, Pierre McKenzie, Dustin Wehr +2
We introduce the Tree Evaluation Problem, show that it is in logDCFL (and hence in P), and study its branching program complexity in the hope of eventually proving a superlogarithm…
Pebbling and Branching Programs Solving the Tree Evaluation Problem
Dustin Wehr
We study restricted computation models related to the Tree Evaluation Problem}. The TEP was introduced in earlier work as a simple candidate for the (*very*) long term goal of sepa…