activity
20242026
collaborators

6 papers

cs.IT2026

Optimal Domain-Aware Privacy Mechanisms for Synthetic Data Generation

Sajani Vithana, Sangwon Jung, Haoyang Hu +3

Differential privacy (DP) imposes fundamental trade-offs between privacy and statistical fidelity in synthetic data generation. While access to public data has been shown to improv…

cs.LG2026

ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport

Atefeh Gilani, Sajani Vithana, Carol Xuan Long +3

Watermarking is an important tool for promoting the responsible use of large language models (LLMs). Existing watermarks insert a signal into generated tokens that either flags LLM…

cs.CR2025

HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions

Dor Tsur, Carol Xuan Long, Claudio Mayrink Verdun +5

Large language model (LLM) watermarks enable authentication of text provenance, curb misuse of machine-generated text, and promote trust in AI systems. Current watermarks operate b…

cs.CV2025

Multi-Group Proportional Representation for Text-to-Image Models

Sangwon Jung, Alex Oesterling, Claudio Mayrink Verdun +3

Text-to-image (T2I) generative models can create vivid, realistic images from textual descriptions. As these models proliferate, they expose new concerns about their ability to rep…

cs.IT2025

Correlated Privacy Mechanisms for Differentially Private Distributed Mean Estimation

Sajani Vithana, Viveck R. Cadambe, Flavio P. Calmon +1

Differentially private distributed mean estimation (DP-DME) is a fundamental building block in privacy-preserving federated learning, where a central server estimates the mean of $…

cs.AI2024

Multi-Group Proportional Representation in Retrieval

Alex Oesterling, Claudio Mayrink Verdun, Carol Xuan Long +5

Image search and retrieval tasks can perpetuate harmful stereotypes, erase cultural identities, and amplify social disparities. Current approaches to mitigate these representationa…