6 papers
World models of environment, agent and joint agent-environment systems
Manuel Baltieri, Filippo Torresan, Yivan Zhang +2
World models are a central component of model-based reinforcement learning. They are usually discussed in terms of what variables they predict, such as observations, rewards, state…
Generalised Bellman recurrence and three dualities in sequential decision-making
Fernando E. Rosas, David Hyland, Daniel Polani
What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through suffic…
Positive Alignment: Artificial Intelligence for Human Flourishing
Ruben Laukkonen, Seb Krier, Chloé Bakalar +13
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psych…
Adaptive state-action abstractions via rate-distortion
Fernando E. Rosas
When learning to walk, infants seem to address a coarse version of the problem first - stay upright, reach the caregiver - and refine it only when further practice at that resoluti…
Symmetries at the origin of hierarchical emergence
Fernando E. Rosas
Many systems of interest exhibit nested emergent layers with their own rules and regularities, and our knowledge about them seems naturally organised around these levels. This pape…
AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability
Fernando Rosas, Alexander Boyd, Manuel Baltieri
Recent work proposes using world models to generate controlled virtual environments in which AI agents can be tested before deployment to ensure their reliability and safety. Howev…