3 papers
cs.LG2026
Chunky Post-Training: Data Driven Failures of Generalization
Seoirse Murray, Allison Qi, Timothy Qian +3
LLM post-training involves many diverse datasets, each targeting a specific behavior. But these datasets encode incidental patterns alongside intended ones: correlations between fo…
cs.AI2025
How Do Large Language Monkeys Get Their Power (Laws)?
Rylan Schaeffer, Joshua Kazdan, John Hughes +7
Recent research across mathematical problem solving, proof assistant programming and multimodal jailbreaking documents a striking finding: when (multimodal) language model tackle a…
cs.CL2024
Best-of-N Jailbreaking
John Hughes, Sara Price, Aengus Lynch +7
We introduce Best-of-N (BoN) Jailbreaking, a simple black-box algorithm that jailbreaks frontier AI systems across modalities. BoN Jailbreaking works by repeatedly sampling variati…