5 papers
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
Johannes von Oswald, Nino Scherrer, Seijin Kobayashi +14
Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention. Although widely adopted, transformers require scaling memory and compu…
Multi-agent cooperation through in-context co-player inference
Marissa A. Weis, Maciej WoÅczyk, Rajai Nasser +4
Achieving cooperation among self-interested agents remains a fundamental challenge in multi-agent reinforcement learning. Recent work showed that mutual cooperation can be induced…
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
Seijin Kobayashi, Yanick Schimpf, Maximilian Schlegel +12
Large-scale autoregressive models pretrained on next-token prediction and finetuned with reinforcement learning (RL) have achieved unprecedented success on many problem domains. Du…
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
Alexander Meulemans, Rajai Nasser, Maciej WoÅczyk +13
The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled from their environment, such that polici…
Multi-agent cooperation through learning-aware policy gradients
Alexander Meulemans, Seijin Kobayashi, Johannes von Oswald +6
Self-interested individuals often fail to cooperate, posing a fundamental challenge for multi-agent learning. How can we achieve cooperation among self-interested, independent lear…