1 paper
Itay Yona, Amir Sarid, Michael Karasik +1
We introduce Doublespeak, a simple in-context representation hijacking attack against large language models (LLMs). The attack works by systematically replacing a harmfu…