1 paper
Yingyun Cui, Yi Xie, Piaohong Wang +3
Coding-agent benchmarks have largely measured whether agents can produce functionally correct patches, but production software also demands measurable speedups on real execution ta…