2 papers
cs.SE2025
SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
Muhammad Shihab Rashid, Christian Bock, Yuan Zhuang +10
Coding agents powered by large language models have shown impressive capabilities in software engineering tasks, but evaluating their performance across diverse programming languag…
cs.SE2024
REDO: Execution-Free Runtime Error Detection for COding Agents
Shou Li, Andrey Kan, Laurent Callot +3
As LLM-based agents exhibit exceptional capabilities in addressing complex problems, there is a growing focus on developing coding agents to tackle increasingly sophisticated tasks…