6 papers
Evaluating Collective Behaviour of Hundreds of LLM Agents
Richard Willis, Jianing Zhao, Yali Du +1
LLM-powered AI assistants acting on behalf of users can produce poor collective outcomes at scale. We introduce a framework for evaluating their emergent behaviour in social dilemm…
SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas
Zihao Guo, Shuqing Shi, Richard Willis +3
Sequential social dilemmas pose a significant challenge in the field of multi-agent reinforcement learning (MARL), requiring environments that accurately reflect the tension betwee…
VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case Study
Zhicheng Zhang, Ziyan Wang, Yali Du +1
Exploration remains a key bottleneck for reinforcement learning (RL) post-training of large language models (LLMs), where sparse feedback and large action spaces can lead to premat…
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
Chandler Smith, Marwa Abdulhai, Manfred Diaz +83
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with bo…
Quantifying the Self-Interest Level of Markov Social Dilemmas
Richard Willis, Yali Du, Joel Z Leibo +1
This paper introduces a novel method for estimating the self-interest level of Markov social dilemmas. We extend the concept of self-interest level from normal-form games to Markov…
Will Systems of LLM Agents Cooperate: An Investigation into a Social Dilemma
Richard Willis, Yali Du, Joel Z Leibo +1
As autonomous agents become more prevalent, understanding their collective behaviour in strategic interactions is crucial. This study investigates the emergent cooperative tendenci…