35 citations · 35 across the 3 of their papers we have counts for
3 papers
cs.AI2026
The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments
Logan Ritchie, Sushant Mehta, Nick Heiner +2
The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments.…
cs.LG2025
FrontierCS: Evolving Challenges for Evolving Intelligence
Qiuyang Mang, Wenhao Chai, Zhifei Li +48
We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competiti…
cs.HC2022★ 35 cited
Measuring Progress on Scalable Oversight for Large Language Models
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez +43
Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on m…