1 paper
Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim +2
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user…