2 citations · 2 across the 8 of their papers we have counts for
Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information
Jiaze Li, Aocheng Shen, Bing Liu +4
Competitive programming is increasingly being used to evaluate the algorithmic reasoning capabilities of large language models (LLMs). However, existing benchmarks primarily focus…
cs.SE2025★ 2 cited
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Xiang Deng, Jeff Da, Edwin Pan +19
We introduce SWE-Bench Pro, a substantially more challenging benchmark that builds upon the best practices of SWE-BENCH [25], but is explicitly designed to capture realistic, compl…