2 papers
cs.LG2024
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
Hannah Brown, Leon Lin, Kenji Kawaguchi +1
We introduce a defense against adversarial attacks on LLMs utilizing self-evaluation. Our method requires no model fine-tuning, instead using pre-trained models to evaluate the inp…
cs.LG2024
Single Character Perturbations Break LLM Alignment
Leon Lin, Hannah Brown, Kenji Kawaguchi +1
When LLMs are deployed in sensitive, human-facing settings, it is crucial that they do not output unsafe, biased, or privacy-violating outputs. For this reason, models are both tra…