7 citations · 9 across the 7 of their papers we have counts for
1 paper · 1 filter
Daniel Zhu, Zihan Wang, Xuchan Bao +1
As language model safeguards become more robust, attackers are pushed toward developing increasingly complex jailbreaks. Prior work has found that this complexity imposes a "jailbr…