1 paper · 1 filter
Sam Herring, Jake Naviasky, Karan Malhotra
Language models are instruction-tuned to refuse harmful requests, but the mechanisms underlying this behavior remain poorly understood. Popular steering methods operate on the resi…