papers

Publications (14)

cs.LG2025

FrontierCS: Evolving Challenges for Evolving Intelligence

Qiuyang Mang, Wenhao Chai, Zhifei Li +48

We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competiti…

cs.CL2022

Discovering Language Model Behaviors with Model-Written Evaluations

Ethan Perez, Sam Ringer, Kamilė Lukošiūtė +60

As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (w…

cs.DS2026

Inverse Quadratic Decay in Random Subset Sum

Edwin Chen, Christof Teuscher

The Subset Sum Problem is a fundamental NP-complete problem in cryptography and combinatorial optimization, with many real-world applications. The Random Subset Sum Problem (RSSP)…

cs.AI2026

EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments

Sushant Mehta, Logan Ritchie, Suhaas Garre +3

We show that training AI agents on high-fidelity reinforcement learning environments produces capabilities that generalize beyond the training distribution. We introduce CoreCraft,…

cs.SE2026

Cross-Benchmark Generalization in Long-Horizon Agents

Sushant Mehta, Logan Ritchie, Liudas Panavas +1

For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templat…

cs.CV2026

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

Suhaas Garre, Emily Ritchie, Sushant Mehta +1

The paper introduces GDP.pdf, a benchmark of professional PDF documents paired with realistic questions to evaluate grounded multimodal reasoning, and reports that current state‑of…

#pdf document understanding#multimodal reasoning#benchmark#layout analysis