1 paper · 1 filter
Xinming Tu, Tianze Wang, Yingzhou +4
As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit ass…