1 paper · 1 filter
Zhiwei Liu, Jielin Qiu, Shiyu Wang +9
The rapid rise of Large Language Models (LLMs)-based intelligent agents underscores the need for robust, scalable evaluation frameworks. Existing methods rely on static benchmarks…