4 papers · 1 filter
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
Spandan Garg, Yufan Huang
While significant progress has been made in automating various aspects of software development through coding agents, there is still significant room for improvement in their bug f…
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
Spandan Garg, Benjamin Steenhoek, Yufan Huang
Current benchmarks for evaluating software engineering agents, such as SWE-Bench Verified, are predominantly derived from GitHub issues and fail to accurately reflect how developer…
PerfBench: Can Agents Resolve Real-World Performance Bugs?
Spandan Garg, Roshanak Zilouchian Moghaddam, Neel Sundaresan
Performance bugs are inefficiencies in software that waste computational resources without causing functional failures, making them particularly challenging to detect and fix. Whil…
Generating Examples From CLI Usage: Can Transformers Help?
Roshanak Zilouchian Moghaddam, Spandan Garg, Colin B. Clement +2
Continuous evolution in modern software often causes documentation, tutorials, and examples to be out of sync with changing interfaces and frameworks. Relying on outdated documenta…