1 paper · 1 filter
Dipesh KC, Anjila Budathoki
Coding-agent benchmarks evaluate whether a single uninterrupted agent can resolve a repository issue. Real software work is messier: tasks are interrupted, reassigned, reviewed, an…