Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Core Safety Values for Provably Corrigible Agents
Aran Nayebi
We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework cons…
cs.AI2025
Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis
Aran Nayebi
We formalize AI alignment as a multi-objective optimization problem called -agreement, in which a set of agents (including humans) must reach…