Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
cs.AI2025
From monoliths to modules: Decomposing transducers for efficient world modelling
Alexander Boyd, Franz Nowak, David Hyland +2
World models have been recently proposed as sandbox environments in which AI agents can be trained and evaluated before deployment. While realistic world models often have high com…
cs.AI2024
Possible Principles for Aligned Structure Learning Agents
Lancelot Da Costa, Tomáš Gavenčiak, David Hyland +5
This paper offers a roadmap for the development of scalable aligned artificial intelligence (AI) from first principle descriptions of natural intelligence. In brief, a possible pat…