#benchmark reliability
try —
2 papers match
cs.SE2026
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Zhi Chen, Zhensu Sun, Yuling Shi +2
The paper audits three repository-level performance‑optimization benchmarks (GSO, SWE‑Perf, SWE‑efficiency) to assess how reliably they measure coding agents, revealing issues with…
#performance optimization#coding agents#benchmark reliability#leaderboard scoring
cs.HC2026
Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment
Ancuta Margondai, Julie Rader, Emma Rader +2
The paper documents how AI systems rapidly surpass human performance on many well-defined tasks, but humans still retain advantages in long‑term reliability, novel problems, and ov…
#ai capability crossing#human‑ai collaboration#cognitive offloading#benchmark reliability