1 paper · 1 filter
Xuan Cuong Ngo, Hao Vo, Ngan Le
Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserv…