1 paper
Sam Herring, Jake Naviasky, Karan Malhotra
Language models are instruction-tuned to refuse harmful requests, but the mechanisms underlying this behavior remain poorly understood. Popular steering methods operate on the resi…