2 papers
cs.LG2026
Jailbroken Frontier Models Retain Their Capabilities
Daniel Zhu, Zihan Wang, Xuchan Bao +1
As language model safeguards become more robust, attackers are pushed toward developing increasingly complex jailbreaks. Prior work has found that this complexity imposes a "jailbr…
cs.AI2026
AI Organizations are More Effective but Less Aligned than Individual Agents
Judy Hanwen Shen, Daniel Zhu, Siddarth Srinivasan +5
AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that multi-agent "AI organizatio…