benchmark reliability 1coding agents 1leaderboard scoring 1performance optimization 1software engineering 1
From the 1 of 3 linked papers with an AI index.
1 citations · 1 across the 3 of their papers we have counts for
Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Zhi Chen, Zhensu Sun, Yuling Shi +2
The paper audits three repository-level performance‑optimization benchmarks (GSO, SWE‑Perf, SWE‑efficiency) to assess how reliably they measure coding agents, revealing issues with…
cs.SE2026
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents
Zhi Chen, Zhensu Sun, Yuling Shi +4
Large Language Model (LLM) code agents increasingly resolve repository-level issues by iteratively editing code, invoking tools, and validating candidate patches. In these workflow…