5 papers
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
Sarthak Choudhary, Atharv Singh Patlan, Nils Palumbo +3
We present Sparse Backdoor, a supply-chain attack that plants a provably undetectable backdoor in pre-trained image classifiers, including convolutional networks and Vision Transfo…
Dependency-Aware Privacy for Multi-turn Agents
Divyam Anshumaan, Sarthak Choudhary, Nils Palumbo +1
LLM agents release private data across multi-service interactions. Existing prompt sanitizers based on metric differential privacy treat each release independently, so adversaries…
SoK: Watermarking for AI-Generated Content
Xuandong Zhao, Sam Gunn, Miranda Christ +11
As the outputs of generative AI (GenAI) techniques improve in quality, it becomes increasingly challenging to distinguish them from human-created content. Watermarking schemes are…
On the Difficulty of Constructing a Robust and Publicly-Detectable Watermark
Jaiden Fairoze, Guillermo Ortiz-Jimenez, Mel Vecerik +2
This work investigates the theoretical boundaries of creating publicly-detectable schemes to enable the provenance of watermarked imagery. Metadata-based approaches like C2PA provi…
Publicly-Detectable Watermarking for Language Models
Jaiden Fairoze, Sanjam Garg, Somesh Jha +3
We present a publicly-detectable watermarking scheme for LMs: the detection algorithm contains no secret information, and it is executable by anyone. We embed a publicly-verifiable…