1 paper · 2 filters
Krishna Kanth Nakka, Xue Jiang, Dmitrii Usynin +1
This paper investigates privacy jailbreaking in LLMs via steering, focusing on whether manipulating activations can bypass LLM alignment and alter response behaviors to privacy rel…