15 papers
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
Mark Russinovich
Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defen…
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Mark Russinovich, Ram Shankar Siva Kumar, Ahmed Salem
Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival record. Assessing this risk via technica…
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
Mark Russinovich, Yanan Cai, Ahmed Salem
Growing concerns over the theft and misuse of Large Language Models (LLMs) underscore the need for effective fingerprinting to link a model to its original version and detect misus…
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
Blake Bullwinkel, Eugenia Kim, Amanda Minnich +1
AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods…
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
Rui Wen, Mark Russinovich, Andrew Paverd +2
Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safety- and privacy-critical appli…
Optimizing Agent Planning for Security and Autonomy
Aashish Kolluri, Rishi Sharma, Manuel Costa +5
Indirect prompt injection attacks threaten AI agents that execute consequential actions, motivating deterministic system-level defenses. Such defenses can provably block unsafe act…