1 paper · 1 filter
Wang Yang, Chaoda Song, Xinpeng Li +7
Existing Agent benchmarks suffer from two critical limitations: high environment interaction overhead (up to 41\% of total evaluation time) and imbalanced task horizon and difficul…