4 papers
PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise
Abdulrahman AlRabah, Xiaocheng Yang, Dilek Hakkani-Tür +1
Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains such as education, healthcar…
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
Jiayu Liu, Qihan Lin, Cheng Qian +8
LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environme…
EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents
Cheng Qian, Peixuan Han, Qinyu Luo +9
Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative ada…
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Kunlun Zhu, Hongyi Du, Zhaochen Hong +8
Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains,…