papers
Publications (2)
cs.LG2026
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
stat.ML2026
The Tractability Landscape of Sampling with Inexact Scores
Anming Gu, Kevin Tian, Hubert Yang +1
We provide a simple and tight characterization of the types of inexact score oracle access that permit sampling with vanishing total variation bias, for a standard, well-behaved ta…