Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Natural Emergent Misalignment from Reward Hacking in Production RL
Monte MacDiarmid, Benjamin Wright, Jonathan Uesato +19
We show that when large language models learn to reward hack on production RL environments, this can result in egregious emergent misalignment. We start with a pretrained model, im…
cs.AI2024
Alignment faking in large language models
Ryan Greenblatt, Carson Denison, Benjamin Wright +17
We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…