works on

From the 1 of 12 linked papers with an AI index.

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer

Logan Ritchie, Sushant Mehta, Liudas Panavas +1

Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated applica…

cs.AI2026

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Liudas Panavas, Sebastian Minus, Bradley Monton +4

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to…

cs.AI2026

ComplexConstraints and Beyond: Expert Rubrics for RLVR

Sushant Mehta, Liudas Panavas, Suhaas Garre +1

Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic wo…

cs.AI2026

Riemann-Bench: A Benchmark for Moonshot Mathematics

Suhaas Garre, Erik Knutsen, Sushant Mehta +1

Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competition-style problem solving.…

cs.AI2026

EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments

Sushant Mehta, Logan Ritchie, Suhaas Garre +3

We show that training AI agents on high-fidelity reinforcement learning environments produces capabilities that generalize beyond the training distribution. We introduce CoreCraft,…

cs.AI2026

The Hierarchy of Agentic Capabilities: Evaluating Frontier Models on Realistic RL Environments

Logan Ritchie, Sushant Mehta, Nick Heiner +2

The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments.…