3 papers
cs.SE2026
AgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent Trajectories
Wuyang Dai, Song Wang
AI coding agents increasingly rely on execution harnesses to interact with repositories and external tools. However, task success does not guarantee reliable execution. Agents may…
cs.SE2026
ABTest: Behavior-Driven Testing for AI Coding Agents
Wuyang Dai, Moses Openja, Hung Viet Pham +3
AI coding agents are increasingly integrated into real-world software development workflows, yet their robustness under diverse and adversarial scenarios remains poorly understood.…
cs.SE2025
BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
Yinghang Ma, Jiho Shin, Leuson Da Silva +5
Recent advances in large language models (LLMs) have accelerated the development of AI-driven automated program repair (APR) solutions. However, these solutions are typically evalu…