33 papers
TabPATE: Differentially Private Tabular In-Context Learning Without Public Data
Dariush Wahdany, Matthew Jagielski, Jesse C. Cresswell +2
Tabular foundation models enable accurate in-context learning (ICL) from small labeled datasets, but the private records placed in context can leak through model predictions. We fi…
Dataset Usage Inference without Shadow Models or Held-out Data
Wojciech Åapacz, StanisÅaw Pawlak, Jan DubiÅski +2
How much of my data was used to train a machine learning model? Dataset Usage Inference (DUI) aims to answer this by estimating what fraction of a dataset contributed to a model's…
Concept Removal for Frontier Image Generative Models
Aditya Kumar, Pierre Joly, Adam Dziedzic +1
Image generative models are trained on massive, largely uncurated internet-scale datasets that contain undesirable visual concepts. Efficiently removing such concepts from the mode…
Natural Identifiers for Privacy and Data Audits in Large Language Models
Lorenzo Rossi, BartÅomiej Marek, Franziska Boenisch +1
Assessing the privacy of large language models (LLMs) presents significant challenges. In particular, most existing methods for auditing differential privacy require the insertion…
MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning
Wenhao Wang, Franziska Boenisch, Michael Backes +1
Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. However, it also causes harmful retention of…
Data Provenance for Image Auto-Regressive Generation
Bihe Zhao, Louis Kerner, Michel Meintz +3
Image autoregressive models (IARs) have recently demonstrated remarkable capabilities in visual content generation, achieving photorealistic quality and rapid synthesis through the…