5 papers
ProtocolBench: Which LLM MultiAgent Protocol to Choose?
Hongyi Du, Jiaqi Su, Jisen Li +6
As large-scale multi-agent systems evolve, the communication protocol layer has become a critical yet under-evaluated factor shaping performance and reliability. Despite the existe…
-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues
Peixuan Han, Hongyi Du, Jiayu Liu +3
Personalization is a crucial capability of modern language agents. However, current research primarily positions personalized agents as passive responders to user preferences, limi…
EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents
Cheng Qian, Peixuan Han, Qinyu Luo +9
Language model agents excel in long-session planning and reasoning, but existing benchmarks primarily focus on goal-oriented tasks with explicit objectives, neglecting creative ada…
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
Cheng Qian, Hongyi Du, Hongru Wang +6
Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect the complexity…
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Kunlun Zhu, Hongyi Du, Zhaochen Hong +8
Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains,…