4 papers
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
Logan Ritchie, Sushant Mehta, Liudas Panavas +1
Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated applica…
Cross-Benchmark Generalization in Long-Horizon Agents
Sushant Mehta, Logan Ritchie, Liudas Panavas +1
For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templat…
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Sushant Mehta, Logan Ritchie, Suhaas Garre +3
We show that training AI agents on high-fidelity reinforcement learning environments produces capabilities that generalize beyond the training distribution. We introduce CoreCraft,…
The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments
Logan Ritchie, Sushant Mehta, Nick Heiner +2
The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments.…