37 citations · 68 across the 11 of their papers we have counts for
13 papers
PerfBench: Can Agents Resolve Real-World Performance Bugs?
Spandan Garg, Roshanak Zilouchian Moghaddam, Neel Sundaresan
Performance bugs are inefficiencies in software that waste computational resources without causing functional failures, making them particularly challenging to detect and fix. Whil…
When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents
Matous Kozak, Roshanak Zilouchian Moghaddam, Siva Sivaraman
LLM-based coding agents are rapidly being deployed in software development, yet their safety implications remain poorly understood. These agents, while capable of accelerating soft…
FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents
Vikram Nitin, Baishakhi Ray, Roshanak Zilouchian Moghaddam
Despite the critical threat posed by software security vulnerabilities, reports are often incomplete, lacking the proof-of-vulnerability (PoV) tests needed to validate fixes and pr…
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
Avi Arora, Jinu Jang, Roshanak Zilouchian Moghaddam
Modern Large Language Model (LLM) agents promise end to end assistance with real-world software tasks, yet existing benchmarks evaluate LLM agents almost exclusively in pre-baked e…
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
Shanchao Liang, Spandan Garg, Roshanak Zilouchian Moghaddam
As large language models (LLMs) become increasingly capable and widely adopted, benchmarks play a central role in assessing their practical utility. For example, SWE-Bench Verified…
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
Dhruv Gautam, Spandan Garg, Jinu Jang +2
Recent advances in language model (LM) agents and function calling have enabled autonomous, feedback-driven systems to solve problems across various digital domains. To better unde…