4 papers
Core Safety Values for Provably Corrigible Agents
Aran Nayebi
We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework cons…
Intrinsic Goals for Autonomous Agents: Model-Based Exploration in Virtual Zebrafish Predicts Ethological Behavior and Whole-Brain Dynamics
Reece Keller, Alyn Kirsch, Felix Pei +3
Autonomy is a hallmark of animal intelligence, enabling adaptive and intelligent behavior in complex environments without relying on external reward or task structure. Existing rei…
Brain-Model Evaluations Need the NeuroAI Turing Test
Jenelle Feather, Meenakshi Khosla, N. Apurva Ratan Murty +1
What makes an artificial system a good model of intelligence? The classical test proposed by Alan Turing focuses on behavior, requiring that an artificial agent's behavior be indis…
Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis
Aran Nayebi
We formalize AI alignment as a multi-objective optimization problem called -agreement, in which a set of agents (including humans) must reach…