Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Distillation Robustifies Unlearning
Bruce W. Lee, Addie Foote, Alex Infanger +6
Current LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: t…
cs.LG2025
Misalignment from Treating Means as Ends
Henrik Marklund, Alex Infanger, Benjamin Van Roy
Reward functions, learned or manually specified, are rarely perfect. Instead of accurately expressing human goals, these reward functions are often distorted by human beliefs about…
cs.LG2024
The Persian Rug: solving toy models of superposition using large-scale symmetries
Aditya Cowsik, Kfir Dolev, Alex Infanger
We present a complete mechanistic description of the algorithm learned by a minimal non-linear sparse data autoencoder in the limit of large input dimension. The model, originally…