3 papers
cs.AI2026
Safety from Honesty in a Disinterested AI Predictor
Yoshua Bengio, Oliver Richardson, Tomáš GavenÄiak +13
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…
cs.AI2025
Possible Principles for Aligned Structure Learning Agents
Lancelot Da Costa, Tomáš GavenÄiak, David Hyland +5
This paper offers a roadmap for the development of scalable aligned artificial intelligence (AI) from first principle descriptions of natural intelligence. In brief, a possible pat…
cs.MA2025
Multi-Agent Risks from Advanced AI
Lewis Hammond, Alan Chan, Jesse Clifton +41
The rapid development of advanced AI agents and the imminent deployment of many instances of these agents will give rise to multi-agent systems of unprecedented complexity. These s…