2 papers
cs.LG2026
The Rollout Infrastructure Tax in Coding-Agent Reinforcement Learning
Daniel Thi Graviet, Lovre Pesut, Ivan Dagelic +2
Coding-agent reinforcement learning treats execution infrastructure as a background implementation detail, despite relying on large numbers of interactive software rollouts. This i…
cs.AI2026
The Capability Frontier: Benchmarks Miss 82% of Model Performance
Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8
Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data…