3 papers
cs.CR2025
Membership Inference Attacks on Sequence Models
Lorenzo Rossi, Michael Aerni, Jie Zhang +1
Sequence models, such as Large Language Models (LLMs) and autoregressive image generators, have a tendency to memorize and inadvertently leak sensitive information. While this tend…
cs.LG2025
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
Kristina Nikolić, Luze Sun, Jie Zhang +1
Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are act…
cs.LG2024
Evaluating the Robustness of the "Ensemble Everything Everywhere" Defense
Jie Zhang, Christian Schlarmann, Kristina Nikolić +4
Ensemble everything everywhere is a defense to adversarial examples that was recently proposed to make image classifiers robust. This defense works by ensembling a model's intermed…