3 papers
cs.LG2026
CheckMIABench: Firm Foundations For Membership Inference Attacks on Language Models
Jeffrey G. Wang, Jason Wang, Marvin Li +1
Membership inference attacks (MIAs) are a canonical way to assess a machine learning model's privacy properties. Although several attempts have been made to evaluate MIAs on langua…
cs.LG2024
Bias Begets Bias: The Impact of Biased Embeddings on Diffusion Models
Sahil Kuchlous, Marvin Li, Jeffrey G. Wang
With the growing adoption of Text-to-Image (TTI) systems, the social biases of these models have come under increased scrutiny. Herein we conduct a systematic investigation of one…
cs.CR2024
Pandora's White-Box: Precise Training Data Detection and Extraction in Large Language Models
Jeffrey G. Wang, Jason Wang, Marvin Li +1
In this paper we develop state-of-the-art privacy attacks against Large Language Models (LLMs), where an adversary with some access to the model tries to learn something about the…