Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems
Ming Shen, Chao Shang, Sadat Shahriar +4
Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: age…
cs.CL2026
Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting
Devang Kulshreshtha, Hang Su, Haibo Jin +2
We introduce \emph{self-jailbreaking}, a threat model in which an aligned LLM guides its own compromise. Unlike most jailbreak techniques, which often rely on handcrafted prompts o…