Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
Jobst Heitzig, Ram Potham
Power is a key concept in AI safety: power-seeking as an instrumental goal, sudden or gradual disempowerment of humans, power balance in human-AI interaction and international AI g…
cs.AI2025
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
Ram Potham, Max Harms
Foundation models (FMs) face a critical safety challenge: as capabilities scale, instrumental convergence drives default trajectories toward loss of human control, potentially culm…