1 paper · 1 filter
Yue Liu, Xiaoxin He, Miao Xiong +5
This paper proposes a simple yet effective jailbreak attack named FlipAttack against black-box LLMs. First, from the autoregressive nature, we reveal that LLMs tend to understand t…