1 paper
Yinger Zhang, Shutong Jiang, Renhao Li +6
While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., tim…