6 papers
Imperfect World Models are Exploitable
Logan Mondal Bhamidipaty, Esmeralda S. Whitammer, David Abel +2
We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred o…
Inter-Agent Relative Representations for Multi-Agent Option Discovery
Raul D. Steleac, Mohan Sridharan, David Abel
Temporally extended actions improve the ability to explore and plan in single-agent settings. In multi-agent settings, the exponential growth of the joint state space with the numb…
Fairness over Equality: Correcting Social Incentives in Asymmetric Sequential Social Dilemmas
Alper Demir, Hüseyin Aydın, Kale-ab Abebe Tessera +2
Sequential Social Dilemmas (SSDs) provide a key framework for studying how cooperation emerges when individual incentives conflict with collective welfare. In Multi-Agent Reinforce…
General agents contain world models
Jonathan Richens, David Abel, Alexis Bellot +1
Are world models a necessary ingredient for flexible, goal-directed behaviour, or is model-free learning sufficient? We provide a formal answer to this question, showing that any a…
Memory Allocation in Resource-Constrained Reinforcement Learning
Massimiliano Tamborski, David Abel
Resource constraints can fundamentally change both learning and decision-making. We explore how memory constraints influence an agent's performance when navigating unknown environm…
A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety
Hyunin Lee, Chanwoo Park, David Abel +1
Black swan events are statistically rare occurrences that carry extremely high risks. A typical view of defining black swan events is heavily assumed to originate from an unpredict…