1 paper
Minyang Hu, Bo Yang, Zhinuo Zhou +4
LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation protocols primarily focus on…