1 paper · 1 filter
Junxi Chen, Junhao Dong, Xiaohua Xie
Recent work has demonstrated the potential of contrastive steering for jailbreaking Large Language Models (LLMs). However, existing methods rely on limited and inherently biased co…