1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.LG2026
Predicting LLM Safety Before Release by Simulating Deployment
Marcus Williams, Hannah Sheahan, Cameron Raymond +8
Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model beha…
cs.LG2024
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Abhay Sheshadri, Aidan Ewart, Phillip Guo +8
Large language models (LLMs) can often be made to behave in undesirable ways that they are explicitly fine-tuned not to. For example, the LLM red-teaming literature has produced a…
cs.CL2024★ 1 cited
Eight Methods to Evaluate Robust Unlearning in LLMs
Aengus Lynch, Phillip Guo, Aidan Ewart +2
Machine unlearning can be useful for removing harmful capabilities and memorized text from large language models (LLMs), but there are not yet standardized methods for rigorously e…