1 paper
Silin Chen, Yufei Yang, Xiaodong Gu +3
Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built upon popular open-sour…