4 papers
CheckMIABench: Firm Foundations For Membership Inference Attacks on Language Models
Jeffrey G. Wang, Jason Wang, Marvin Li +1
Membership inference attacks (MIAs) are a canonical way to assess a machine learning model's privacy properties. Although several attempts have been made to evaluate MIAs on langua…
Bias Begets Bias: The Impact of Biased Embeddings on Diffusion Models
Sahil Kuchlous, Marvin Li, Jeffrey G. Wang
With the growing adoption of Text-to-Image (TTI) systems, the social biases of these models have come under increased scrutiny. Herein we conduct a systematic investigation of one…
Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models
Jeffrey G. Wang, Jason Wang, Marvin Li +1
In this paper we develop state-of-the-art privacy attacks against Large Language Models (LLMs), where an adversary with some access to the model tries to learn something about the…
MoPe: Model Perturbation-based Privacy Attacks on Language Models
Marvin Li, Jason Wang, Jeffrey Wang +1
Recent work has shown that Large Language Models (LLMs) can unintentionally leak sensitive information present in their training data. In this paper, we present Model Perturbations…