1 paper · 1 filter
Yuling Shi, Zhensu Sun, Junsen Dong +3
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hun…