1 citations · 1 across the 1 of their papers we have counts for
1 paper
Wenqi Huang, Charley Lee, Leonard Tng +1
DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks follow SWE-bench in mining merge…