8 citations · 8 across the 3 of their papers we have counts for
1 paper · 1 filter
Jingbo Zhou, Yusai Zhao, Qi Bao +12
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents…