5 papers · 1 filter
A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences
Jobst Heitzig, Ram Potham
This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between h…
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Dewi Gould, Francis Rhys Ward, Anders Cairns Woodruff +18
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning intern…
Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety
Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad +3
An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but unt…
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
Jobst Heitzig, Ram Potham
Power is a key concept in AI safety: power-seeking as an instrumental goal, sudden or gradual disempowerment of humans, power balance in human-AI interaction and international AI g…
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
Ram Potham, Max Harms
Foundation models (FMs) face a critical safety challenge: as capabilities scale, instrumental convergence drives default trajectories toward loss of human control, potentially culm…