4 papers
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
Jobst Heitzig, Ram Potham
Power is a key concept in AI safety: power-seeking as an instrumental goal, sudden or gradual disempowerment of humans, power balance in human-AI interaction and international AI g…
MAEBE: Multi-Agent Emergent Behavior Framework
Sinem Erisken, Timothy Gothard, Martin Leitgab +1
Traditional AI safety evaluations on isolated LLMs are insufficient as multi-agent AI ensembles become prevalent, introducing novel emergent risks. This paper introduces the Multi-…
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
Ram Potham
Credible safety plans for advanced AI development require methods to verify agent behavior and detect potential control deficiencies early. A fundamental aspect is ensuring agents…
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
Ram Potham, Max Harms
Foundation models (FMs) face a critical safety challenge: as capabilities scale, instrumental convergence drives default trajectories toward loss of human control, potentially culm…