Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Imperfect World Models are Exploitable
Logan Mondal Bhamidipaty, Esmeralda S. Whitammer, David Abel +2
We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred o…
cs.AI2025
General agents contain world models
Jonathan Richens, David Abel, Alexis Bellot +1
Are world models a necessary ingredient for flexible, goal-directed behaviour, or is model-free learning sufficient? We provide a formal answer to this question, showing that any a…
cs.AI2025
A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety
Hyunin Lee, Chanwoo Park, David Abel +1
Black swan events are statistically rare occurrences that carry extremely high risks. A typical view of defining black swan events is heavily assumed to originate from an unpredict…