2 papers
cs.SE2026
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns
Vlad Sobal, Shuo Yang, Yuting Zhang +2
We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests) they can handle before failing. Unlike t…
cs.LG2026
EvoMAS: Evolutionary Generation of Multi-Agent Systems
Yuntong Hu, Yuting Zhang, Matthew Trager +4
Large language model (LLM)-based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but designing effective MAS architectures…