1 paper
Xinming Tu, Tianze Wang, Yingzhou +4
As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit ass…