1 paper · 1 filter
Suliu Qin, Lu Yin, Xilu Wang
Language-model agents increasingly tackle long-horizon tasks in interactive environments, yet their evaluation commonly relies on task-level success rates by reducing an entire exe…