11 papers
CryptanalysisBench: Can LLMs do Cryptanalysis?
Lukas Fluri, Avital Shafran, Nicholas Carlini +5
Cryptanalysis - the task of finding attacks against cryptographic schemes - sits at the intersection of mathematical reasoning and cybersecurity, two areas where LLMs have advanced…
Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration
Debeshee Das, Julien Piet, Darya Kaviani +3
Memory systems enable otherwise-stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. We characterize the Trojan Hippo attack,…
Large-scale online deanonymization with LLMs
Simon Lermen, Daniel Paleka, Joshua Swanson +3
We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer…
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Milad Nasr, Nicholas Carlini, Chawin Sitawarin +11
How should we evaluate the robustness of language model defenses? Current defenses against jailbreaks and prompt injections (which aim to prevent an attacker from eliciting harmful…
SoK: Watermarking for AI-Generated Content
Xuandong Zhao, Sam Gunn, Miranda Christ +11
As the outputs of generative AI (GenAI) techniques improve in quality, it becomes increasingly challenging to distinguish them from human-created content. Watermarking schemes are…
International Scientific Report on the Safety of Advanced AI (Interim Report)
Yoshua Bengio, Sören Mindermann, Daniel Privitera +41
This is the interim publication of the first International Scientific Report on the Safety of Advanced AI. The report synthesises the scientific understanding of general-purpose AI…