activity
20192025
most citedOn the Efficacy of Differentially Private Few-shot Image Classification

4 citations · 4 across the 6 of their papers we have counts for

collaborators

8 papers

cs.CR2025

Beyond Membership: Limitations of Add/Remove Adjacency in Differential Privacy

Gauri Pradhan, Joonas Jälkö, Santiago Zanella-Béguelin +1

Training machine learning models with differential privacy (DP) limits an adversary's ability to infer sensitive information about the training data. It can be interpreted as a bou…

cs.CR2025

A Systematization of Security Vulnerabilities in Computer Use Agents

Daniel Jones, Giorgio Severi, Martin Pouliot +7

Computer Use Agents (CUAs), autonomous systems that interact with software interfaces via browsers or virtual machines, are rapidly being deployed in consumer and enterprise enviro…

cs.CR2025

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks

Blake Bullwinkel, Mark Russinovich, Ahmed Salem +8

Recent research has demonstrated that state-of-the-art LLMs and defenses remain susceptible to multi-turn jailbreak attacks. These attacks require only closed-box model access and…

cs.LG2024

Permissive Information-Flow Analysis for Large Language Models

Shoaib Ahmed Siddiqui, Radhika Gaonkar, Boris Köpf +7

Large Language Models (LLMs) are rapidly becoming commodity components of larger software systems. This poses natural security and privacy problems: poisoned data retrieved from on…

cs.CR2024

Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition

Edoardo Debenedetti, Javier Rando, Daniel Paleka +18

Large language model systems face important security risks from maliciously crafted messages that aim to overwrite the system's original instructions or leak private data. To study…

cs.CR2024

Closed-Form Bounds for DP-SGD against Record-level Inference

Giovanni Cherubin, Boris Köpf, Andrew Paverd +3

Machine learning models trained with differentially-private (DP) algorithms such as DP-SGD enjoy resilience against a wide range of privacy attacks. Although it is possible to deri…