1 paper
Alejandro Cuadron, Pengfei Yu, Yang Liu +1
Despite rapid progress in LLM agents, performance on long-horizon, tool-using tasks remains fragile. To better understand this fragility, we ask a simple question: \emph{do all act…