Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
The Persistent Vulnerability of Aligned AI Systems
Aengus Lynch
Autonomous AI agents are being deployed with filesystem access, email control, and multi-step planning. This thesis contributes to four open problems in AI safety: understanding da…
cs.LG2025
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Abhay Sheshadri, Aidan Ewart, Phillip Guo +8
Large language models (LLMs) can often be made to behave in undesirable ways that they are explicitly fine-tuned not to. For example, the LLM red-teaming literature has produced a…
cs.LG2025
Analyzing the Generalization and Reliability of Steering Vectors
Daniel Tan, David Chanin, Aengus Lynch +4
Steering vectors (SVs) have been proposed as an effective approach to adjust language model behaviour at inference time by intervening on intermediate model activations. They have…