2 papers
cs.CR2025
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
Stuart Armstrong, Matija Franklin, Connor Stevens +1
Recent work showed Best-of-N (BoN) jailbreaking using repeated use of random augmentations (such as capitalization, punctuation, etc) is effective against all major large language…
cs.AI2023
CoinRun: Solving Goal Misgeneralisation
Stuart Armstrong, Alexandre Maranhão, Oliver Daniels-Koch +2
Goal misgeneralisation is a key challenge in AI alignment -- the task of getting powerful Artificial Intelligences to align their goals with human intentions and human morality. In…