2 papers
cs.CL2025
Endless Jailbreaks with Bijection Learning
Brian R. Y. Huang, Maximilian Li, Leonard Tang
Despite extensive safety measures, LLMs are vulnerable to adversarial inputs, or jailbreaks, which can elicit unsafe behaviors. In this work, we introduce bijection learning, a pow…
cs.CL2024
Plentiful Jailbreaks with String Compositions
Brian R. Y. Huang
Large language models (LLMs) remain vulnerable to a slew of adversarial attacks and jailbreaking methods. One common approach employed by white-hat attackers, or red-teamers, is to…