1 paper · 1 filter
Moreno D'IncÃ, Nicu Sebe, Massimiliano Mancini
To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feas…