2 citations · 3 across the 14 of their papers we have counts for
1 paper · 1 filter
Chao Peng, Zhiheng Lyu, Peijie Dong +2
Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. Mo…