3 papers
cs.AI2026
Consequentialist Objectives and Catastrophe
Henrik Marklund, Alex Infanger, Benjamin Van Roy
Because human preferences are too complex to codify, AIs operate with misspecified objectives. Optimizing such objectives often produces undesirable outcomes; this phenomenon is kn…
cs.LG2025
Distillation Robustifies Unlearning
Bruce W. Lee, Addie Foote, Alex Infanger +6
Current LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: t…
cs.LG2025
Misalignment from Treating Means as Ends
Henrik Marklund, Alex Infanger, Benjamin Van Roy
Reward functions, learned or manually specified, are rarely perfect. Instead of accurately expressing human goals, these reward functions are often distorted by human beliefs about…