3 papers
cs.LG2026
A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
Hong Kiat Tan, Linh Le, David Williams-King
Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. The evaluations that guide frontier-model deployment and governance th…
cs.LG2026
Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities
Linh Le, Hong Kiat Tan, David Williams-King
Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governan…
cs.LG2026
Efficient Safety Alignment of Language Models via Latent Personality Traits
Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere +3
Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT)…