1 paper · 1 filter
Thomas Winninger, Boussad Addad, Katarzyna Kapusta
Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechani…