2 papers
cs.LG2026
Stress-Testing Alignment Audits With Prompt-Level Strategic Deception
Oliver Daniels, Perusha Moodley, Benjamin M. Marlin +1
Alignment audits aim to robustly identify hidden goals from strategic, situationally aware misaligned models. Despite this threat model, existing auditing methods have not been sys…
cs.LG2025
ACE and Diverse Generalization via Selective Disagreement
Oliver Daniels, Stuart Armstrong, Alexandre Maranhão +3
Deep neural networks are notoriously sensitive to spurious correlations - where a model learns a shortcut that fails out-of-distribution. Existing work on spurious correlations has…