4 papers · 1 filter
A game theory for foundation models shows new paths to rational cooperation through similarity inference
Alexander Meulemans, Maciej Wołczyk, Maciej WoÅczyk +14
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is…
Multi-agent cooperation through in-context co-player inference
Marissa A. Weis, Maciej WoÅczyk, Rajai Nasser +4
Achieving cooperation among self-interested agents remains a fundamental challenge in multi-agent reinforcement learning. Recent work showed that mutual cooperation can be induced…
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
Alexander Meulemans, Rajai Nasser, Maciej WoÅczyk +13
The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled from their environment, such that polici…
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
Scott Emmons, Erik Jenner, David K. Elson +5
While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on "unfaithfulness" has cast doubt on its reliability. These findings highlight an important…