1 citations · 1 across the 3 of their papers we have counts for
9 papers
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Mark Russinovich, Ram Shankar Siva Kumar, Ahmed Salem
Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival record. Assessing this risk via technica…
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
Mark Russinovich, Yanan Cai, Ahmed Salem
Growing concerns over the theft and misuse of Large Language Models (LLMs) underscore the need for effective fingerprinting to link a model to its original version and detect misus…
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
Rui Wen, Mark Russinovich, Andrew Paverd +2
Backdoor attacks pose a serious security threat to large language models (LLMs), which are increasingly deployed as general-purpose assistants in safety- and privacy-critical appli…
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
Mark Russinovich, Yanan Cai, Keegan Hines +3
Safety alignment is only as robust as its weakest failure mode. Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-…
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
Blake Bullwinkel, Mark Russinovich, Ahmed Salem +8
Recent research has demonstrated that state-of-the-art LLMs and defenses remain susceptible to multi-turn jailbreak attacks. These attacks require only closed-box model access and…
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
Yanan Cai, Ahmed Salem, Besmira Nushi +1
We introduce LogiPlan, a novel benchmark designed to evaluate the capabilities of large language models (LLMs) in logical planning and reasoning over complex relational structures.…