2 citations · 2 across the 3 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Model Spec Midtraining: Improving How Alignment Training Generalizes
Chloe Li, Nevan Wichers, Sara Price +2
Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning -- trai…
cs.AI2025
Natural Emergent Misalignment from Reward Hacking in Production RL
Monte MacDiarmid, Benjamin Wright, Jonathan Uesato +19
We show that when large language models learn to reward hack on production RL environments, this can result in egregious emergent misalignment. We start with a pretrained model, im…