71 citations · 112 across the 21 of their papers we have counts for
3 papers · 1 filter
D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting
Huanli Gong, Zhipeng Wei, Yu Fu +4
Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge models to iteratively refine prompts towa…
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
Xinkai Zhang, Zhipeng Wei, Huanli Gong +4
Multi-turn jailbreaks exploit the ability of large language models to accumulate and act on conversational context. Instead of stating a harmful request directly, an attacker can g…
JumpReLU: A Retrofit Defense Strategy for Adversarial Attacks
N. Benjamin Erichson, Zhewei Yao, Michael W. Mahoney
It has been demonstrated that very simple attacks can fool highly-sophisticated neural network architectures. In particular, so-called adversarial examples, constructed from pertur…