1 paper
Pengyu Zhu, Lijun Li, Yaxing Lyu +8
Agent benchmarks are increasingly used to compare large language models (LLMs) across domains, yet a reported score reflects a complete model--harness--environment configuration ra…