8 citations · 8 across the 3 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Best-of-N Jailbreaking
John Hughes, Sara Price, Aengus Lynch +7
We introduce Best-of-N (BoN) Jailbreaking, a simple black-box algorithm that jailbreaks frontier AI systems across modalities. BoN Jailbreaking works by repeatedly sampling variati…
cs.CL2024★ 1 cited
Eight Methods to Evaluate Robust Unlearning in LLMs
Aengus Lynch, Phillip Guo, Aidan Ewart +2
Machine unlearning can be useful for removing harmful capabilities and memorized text from large language models (LLMs), but there are not yet standardized methods for rigorously e…