1 paper · 1 filter
Jonathan Hayase, Ema Borevkovic, Nicholas Carlini +2
Recent work has shown it is possible to construct adversarial examples that cause an aligned language model to emit harmful strings or perform harmful behavior. Existing attacks wo…