25 citations · 27 across the 4 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Jailbroken Frontier Models Retain Their Capabilities
Daniel Zhu, Zihan Wang, Xuchan Bao +1
As language model safeguards become more robust, attackers are pushed toward developing increasingly complex jailbreaks. Prior work has found that this complexity imposes a "jailbr…
cs.LG2025
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
Narmeen Oozeer, Luke Marks, Shreyans Jain +2
Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of…