7 papers
PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios
Chris Zhu, Sasha Cui, Will Sanok Dufallo +4
We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and economic value creation. To do so, we in…
Lemon Agent Technical Report
Haipeng Jiang, Kailong Ren, Zimo Yin +17
Recent advanced LLM-powered agent systems have exhibited their remarkable capabilities in tackling complex, long-horizon tasks. Nevertheless, they still suffer from inherent limita…
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
Zhiyuan Hu, Yunhai Hu, Juncheng Liu +9
Multi-agent systems have evolved into practical LLM-driven collaborators for many applications, gaining robustness from diversity and cross-checking. However, multi-agent RL (MARL)…
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
Junhao Liu, Zhenhao Xu, Yuxin Fang +3
Recently, there have been notable advancements in large language models (LLMs), demonstrating their growing abilities in complex reasoning. However, existing research largely overl…
Towards impactful challenges: post-challenge paper, benchmarks and other dissemination actions
Antoine Marot, David Rousseau, Zhen +1
The conclusion of an AI challenge is not the end of its lifecycle; ensuring a long-lasting impact requires meticulous post-challenge activities. The long-lasting impact also needs…
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
Pengfei Yu, Dongming Shen, Silin Meng +8
We present RPGBench, the first benchmark designed to evaluate large language models (LLMs) as text-based role-playing game (RPG) engines. RPGBench comprises two core tasks: Game Cr…