1 paper · 1 filter
Ely Hahami, Ishaan Sinha, Lavik Jain
Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering: injecting concept…