1 paper
Jiajun Jiang, Sharon Zheng, Natan Vidra +1
AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to the number of unit tests in a single sub…