11 citations · 12 across the 4 of their papers we have counts for
Showing cs.MAShow all
2 papers · 1 filter
cs.MA2026
SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
Grant Hamblin, Kevin Song, Zhanda Zhu +4
Software engineering (SWE) agents are transitioning from code generation to full software development lifecycle automation. A critical phase in this lifecycle is specification desi…
cs.MA2025★ 1 cited
Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents
Kevin Song, Anand Jayarajan, Yaoyao Ding +4
Large Language Models (LLMs) agents augmented with domain tools promise to autonomously execute complex tasks requiring human-level intelligence, such as customer service and digit…