5 papers
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
Mark Russinovich, Yanan Cai, Keegan Hines +3
Safety alignment is only as robust as its weakest failure mode. Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-…
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
Blake Bullwinkel, Mark Russinovich, Ahmed Salem +8
Recent research has demonstrated that state-of-the-art LLMs and defenses remain susceptible to multi-turn jailbreak attacks. These attacks require only closed-box model access and…
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
Yanan Cai, Ahmed Salem, Besmira Nushi +1
We introduce LogiPlan, a novel benchmark designed to evaluate the capabilities of large language models (LLMs) in logical planning and reasoning over complex relational structures.…
Jailbreaking is (Mostly) Simpler Than You Think
Mark Russinovich, Ahmed Salem
We introduce the Context Compliance Attack (CCA), a novel, optimization-free method for bypassing AI safety mechanisms. Unlike current approaches -- which rely on complex prompt en…
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
Mark Russinovich, Ahmed Salem
Recent copyright agreements between AI companies and content creators underscore the need for fine-grained control over language models' ability to reproduce copyrighted text. Exis…