1 paper
Philipp E. Glass, Allan Tucker, Yongmin Li +1
Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to…