Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
Logan Ritchie, Sushant Mehta, Liudas Panavas +1
Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated applica…
cs.AI2026
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Liudas Panavas, Sebastian Minus, Bradley Monton +4
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to…
cs.AI2026
ComplexConstraints and Beyond: Expert Rubrics for RLVR
Sushant Mehta, Liudas Panavas, Suhaas Garre +1
Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic wo…