1 paper · 1 filter
Samuel Soo, Chen Guang, Wesley Teng +3
Effective and reliable control over large language model (LLM) behavior is a significant challenge. While activation steering methods, which add steering vectors to a model's hidde…