1 paper
Jie Wu, Ming Gong, Feixiang Cheng +1
Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite…