collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

Jobst Heitzig, Ram Potham

This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between h…

cs.AI2026

Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

Dewi Gould, Francis Rhys Ward, Anders Cairns Woodruff +18

Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perform sufficiently complex reasoning intern…

cs.AI2026

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad +3

An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but unt…

cs.AI2025

Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power

Jobst Heitzig, Ram Potham

Power is a key concept in AI safety: power-seeking as an instrumental goal, sudden or gradual disempowerment of humans, power balance in human-AI interaction and international AI g…

cs.AI2025

Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models

Ram Potham, Max Harms

Foundation models (FMs) face a critical safety challenge: as capabilities scale, instrumental convergence drives default trajectories toward loss of human control, potentially culm…