100 citations · 381 across the 48 of their papers we have counts for
1 paper · 2 filters
Jonathan Hayase, Ema Borevkovic, Nicholas Carlini +2
Recent work has shown it is possible to construct adversarial examples that cause an aligned language model to emit harmful strings or perform harmful behavior. Existing attacks wo…