4 papers
(Mis)generalization of Helpful-only Fine-tuning
Mohammad Omar Khursheed, Baram Sosis, Fabien Roger
Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of AI R&D where refusals would…
Steering Language Models with Weight Arithmetic
Constanza Fierro, Fabien Roger
Providing high-quality feedback to Large Language Models (LLMs) on a diverse training distribution can be difficult and expensive, and providing feedback only on a narrow distribut…
Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation
Callum Canavan, Aditya Shrivastava, Allison Qi +2
To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones…
Do Unlearning Methods Remove Information from Language Model Weights?
Aghyad Deeb, Fabien Roger
Large Language Models' knowledge of how to perform cyber-security attacks, create bioweapons, and manipulate humans poses risks of misuse. Previous work has proposed methods to unl…