1 paper
Preethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde +2
Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. T…