4 papers · 1 filter
World models of environment, agent and joint agent-environment systems
Manuel Baltieri, Filippo Torresan, Yivan Zhang +2
World models are a central component of model-based reinforcement learning. They are usually discussed in terms of what variables they predict, such as observations, rewards, state…
Positive Alignment: Artificial Intelligence for Human Flourishing
Ruben Laukkonen, Seb Krier, Chloé Bakalar +13
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psych…
From monoliths to modules: Decomposing transducers for efficient world modelling
Alexander Boyd, Franz Nowak, David Hyland +2
World models have been recently proposed as sandbox environments in which AI agents can be trained and evaluated before deployment. While realistic world models often have high com…
AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability
Fernando Rosas, Alexander Boyd, Manuel Baltieri
Recent work proposes using world models to generate controlled virtual environments in which AI agents can be tested before deployment to ensure their reliability and safety. Howev…