Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
cs.AI2026
From monoliths to modules: Decomposing transducers for efficient world modelling
Alexander Boyd, Franz Nowak, David Hyland +2
World models have been recently proposed as sandbox environments in which AI agents can be trained and evaluated before deployment. While realistic world models often have high com…
cs.AI2025
Possible Principles for Aligned Structure Learning Agents
Lancelot Da Costa, Tomáš GavenÄiak, David Hyland +5
This paper offers a roadmap for the development of scalable aligned artificial intelligence (AI) from first principle descriptions of natural intelligence. In brief, a possible pat…