1 paper · 1 filter
Diaoulé Diallo, Katharina Dworatzyk, Sophie Jentzsch +3
Controlling the behavior of large language models (LLMs) at inference time is essential for aligning outputs with human abilities and safety requirements. \emph{Activation steering…