6 papers
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
Carissa Cullen, Harry Garland, Alexander Roman +3
Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectories. The Discounted Reward for Sa…
Shutdownable Agents through POST-Agency
Elliott Thornley
Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happen. I propose that we train agents to sat…
Out-of-Distribution Generalization of Risk Aversion in Language Models
Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie +3
Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-rewa…
Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training
Christian Moya, Alex Semendinger, Guang Lin +1
Preference learning methods like Direct Preference Optimization (DPO) are known to induce reliance on spurious correlations, leading to sycophancy and length bias in today's langua…
Towards Shutdownable Agents via Stochastic Choice
Elliott Thornley, Alexander Roman, Christos Ziakas +2
The POST-Agents Proposal (PAP) is an idea for ensuring that advanced artificial agents never resist shutdown. A key part of the PAP is using a novel `Discounted Reward for Same-Len…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…