6 papers
Out-of-Distribution Generalization of Risk Aversion in Language Models
Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie +3
Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-rewa…
Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training
Christian Moya, Alex Semendinger, Guang Lin +1
Preference learning methods like Direct Preference Optimization (DPO) are known to induce reliance on spurious correlations, leading to sycophancy and length bias in today's langua…
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
Carissa Cullen, Harry Garland, Alexander Roman +3
Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length trajectories. The Discounted Reward for Sa…
Shutdownable Agents through POST-Agency
Elliott Thornley
Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happen. I propose that we train agents to sat…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Towards Shutdownable Agents via Stochastic Choice
Elliott Thornley, Alexander Roman, Christos Ziakas +2
The POST-Agents Proposal (PAP) is an idea for ensuring that advanced artificial agents never resist shutdown. A key part of the PAP is using a novel `Discounted Reward for Same-Len…