computer-use benchmarking 1interactive agents 1long-horizon tasks 1safety auditing 1tool-use evaluation 1
From the 1 of 12 linked papers with an AI index.
Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
Gabriel Orlanski, Devjeet Roy, Alexander Yun +7
Software development is iterative, yet agentic coding benchmarks hide design issues through their single-shot setup. Recent iterative benchmarks attempt to remedy this but heavily…
cs.SE2026
Pareto Optimal Code Generation
Gabriel Orlanski, Nicholas Roberts, Aws Albarghouthi +1
Generate-then-rank is the dominant test-time scaling (TTS) paradigm for code generation, but scaling accuracy by sampling and executing more candidates makes comprehensive verifica…