From the 1 of 14 linked papers with an AI index.
14 papers
Optimal Domain-Aware Privacy Mechanisms for Synthetic Data Generation
Sajani Vithana, Sangwon Jung, Haoyang Hu +3
Differential privacy (DP) imposes fundamental trade-offs between privacy and statistical fidelity in synthetic data generation. While access to public data has been shown to improv…
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories
Roy Rinberg, Usha Bhalla, Igor Shilov +2
The paper presents RippleBench, a benchmark that automatically creates multiple‑choice questions about semantically related concepts using a Wikipedia‑based retrieval system, to me…
Sort, Partition, Randomize: Optimal Binary Hypothesis Testing under Local Differential Privacy
Elena Ghazi, Jawad Nasser, Flavio Calmon +1
We study optimal design of -locally differentially private mechanisms for binary hypothesis testing. Each observation is drawn from one of two known distributions $P_0…
ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
Atefeh Gilani, Sajani Vithana, Carol Xuan Long +3
Watermarking is an important tool for promoting the responsible use of large language models (LLMs). Existing watermarks insert a signal into generated tokens that either flags LLM…
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
Usha Bhalla, Alex Oesterling, Claudio Mayrink Verdun +2
Translating the internal representations and computations of models into concepts that humans can understand is a key goal of interpretability. While recent dictionary learning met…
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
Dor Tsur, Carol Xuan Long, Claudio Mayrink Verdun +5
Large language model (LLM) watermarks enable authentication of text provenance, curb misuse of machine-generated text, and promote trust in AI systems. Current watermarks operate b…