1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Thomas Winninger, Boussad Addad, Katarzyna Kapusta
Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechani…