2 papers
cs.CY2026
A Virtuous AI is an Existential Risk
Guillermo Del Pinal, Youngchan Lee, Min Ohn
This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) on…
cs.AI2026
Emergent alignment and the projectability of ethical personas
Guillermo Del Pinal, Youngchan Lee, Calum McNamara +1
Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection' (PSM) hypothesis: during pr…