6 papers
Optimal Domain-Aware Privacy Mechanisms for Synthetic Data Generation
Sajani Vithana, Sangwon Jung, Haoyang Hu +3
Differential privacy (DP) imposes fundamental trade-offs between privacy and statistical fidelity in synthetic data generation. While access to public data has been shown to improv…
ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
Atefeh Gilani, Sajani Vithana, Carol Xuan Long +3
Watermarking is an important tool for promoting the responsible use of large language models (LLMs). Existing watermarks insert a signal into generated tokens that either flags LLM…
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
Dor Tsur, Carol Xuan Long, Claudio Mayrink Verdun +5
Large language model (LLM) watermarks enable authentication of text provenance, curb misuse of machine-generated text, and promote trust in AI systems. Current watermarks operate b…
Multi-Group Proportional Representation for Text-to-Image Models
Sangwon Jung, Alex Oesterling, Claudio Mayrink Verdun +3
Text-to-image (T2I) generative models can create vivid, realistic images from textual descriptions. As these models proliferate, they expose new concerns about their ability to rep…
Correlated Privacy Mechanisms for Differentially Private Distributed Mean Estimation
Sajani Vithana, Viveck R. Cadambe, Flavio P. Calmon +1
Differentially private distributed mean estimation (DP-DME) is a fundamental building block in privacy-preserving federated learning, where a central server estimates the mean of $…
Multi-Group Proportional Representation in Retrieval
Alex Oesterling, Claudio Mayrink Verdun, Carol Xuan Long +5
Image search and retrieval tasks can perpetuate harmful stereotypes, erase cultural identities, and amplify social disparities. Current approaches to mitigate these representationa…