Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Scaling Trends in Language Model Robustness
Nikolaus Howe, Ian McKenzie, Oskar Hollinsworth +5
Increasing model size has unlocked a dazzling array of capabilities in modern language models. At the same time, even frontier models remain vulnerable to jailbreaks and prompt inj…
cs.LG2025
Can Go AIs be adversarially robust?
Tom Tseng, Euan McLean, Kellin Pelrine +2
Prior work found that superhuman Go AIs can be defeated by simple adversarial strategies, especially "cyclic" attacks. In this paper, we study whether adding natural countermeasure…