4 papers · 1 filter
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
Carissa Cullen, Harry Garland, Alexander Roman +3
Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectories. The Discounted Reward for Sa…
Shutdownable Agents through POST-Agency
Elliott Thornley
Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happen. I propose that we train agents to sat…
Towards Shutdownable Agents via Stochastic Choice
Elliott Thornley, Alexander Roman, Christos Ziakas +2
The POST-Agents Proposal (PAP) is an idea for ensuring that advanced artificial agents never resist shutdown. A key part of the PAP is using a novel `Discounted Reward for Same-Len…
The Shutdown Problem: An AI Engineering Puzzle for Decision Theorists
Elliott Thornley
I explain the shutdown problem: the problem of designing artificial agents that (1) shut down when a shutdown button is pressed, (2) don't try to prevent or cause the pressing of t…