3 papers
cs.LG2025
ACE and Diverse Generalization via Selective Disagreement
Oliver Daniels, Stuart Armstrong, Alexandre Maranhão +3
Deep neural networks are notoriously sensitive to spurious correlations - where a model learns a shortcut that fails out-of-distribution. Existing work on spurious correlations has…
cs.CR2025
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
Stuart Armstrong, Matija Franklin, Connor Stevens +1
Recent work showed Best-of-N (BoN) jailbreaking using repeated use of random augmentations (such as capitalization, punctuation, etc) is effective against all major large language…
cs.AI2022
The dangers in algorithms learning humans' values and irrationalities
Rebecca Gorman, Stuart Armstrong
For an artificial intelligence (AI) to be aligned with human values (or human preferences), it must first learn those values. AI systems that are trained on human behavior, risk mi…