Publications (14)
FrontierCS: Evolving Challenges for Evolving Intelligence
Qiuyang Mang, Wenhao Chai, Zhifei Li +48
We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competiti…
Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez, Sam Ringer, KamilÄ LukoÅ¡iÅ«tÄ +60
As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (w…
Inverse Quadratic Decay in Random Subset Sum
Edwin Chen, Christof Teuscher
The Subset Sum Problem is a fundamental NP-complete problem in cryptography and combinatorial optimization, with many real-world applications. The Random Subset Sum Problem (RSSP)…
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Sushant Mehta, Logan Ritchie, Suhaas Garre +3
We show that training AI agents on high-fidelity reinforcement learning environments produces capabilities that generalize beyond the training distribution. We introduce CoreCraft,…
Cross-Benchmark Generalization in Long-Horizon Agents
Sushant Mehta, Logan Ritchie, Liudas Panavas +1
For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templat…
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
Suhaas Garre, Emily Ritchie, Sushant Mehta +1
The paper introduces GDP.pdf, a benchmark of professional PDF documents paired with realistic questions to evaluate grounded multimodal reasoning, and reports that current state‑of…