6 papers
PACZero: PAC-Private Fine-Tuning of Language Models via Sign Quantization
Murat Bilgehan Ertan, Xiaochen Zhu, Phuong Ha Nguyen +2
We introduce PACZero, a family of PAC-private zeroth-order mechanisms for fine-tuning large language models that delivers usable utility at . This privacy regime…
Trade-off Functions for DP-SGD with Subsampling based on Random Shuffling: Tight Upper and Lower Bounds
Marten van Dijk, Murat Bilgehan Ertan
We derive a tight analysis of the trade-off function for Differentially Private Stochastic Gradient Descent (DP-SGD) with subsampling based on random shuffling within the -DP fr…
Fundamental Limitations of Favorable Privacy-Utility Guarantees for DP-SGD
Murat Bilgehan Ertan, Marten van Dijk
Differentially Private Stochastic Gradient Descent (DP-SGD) is the dominant paradigm for private training, but its fundamental limitations under worst-case adversarial privacy defi…
TOSSS: a CVE-based Software Security Benchmark for Large Language Models
Marc Damie, Murat Bilgehan Ertan, Domenico Essoussi +3
With their increasing capabilities, Large Language Models (LLMs) are now used across many industries. They have become useful tools for software engineers and support a wide range…
On the Evidentiary Limits of Membership Inference for Copyright Auditing
Murat Bilgehan Ertan, Emirhan Böge, Min Chen +2
As large language models (LLMs) are trained on increasingly opaque corpora, membership inference attacks (MIAs) have been proposed to audit whether copyrighted texts were used duri…
Beyond Anonymization: Object Scrubbing for Privacy-Preserving 2D and 3D Vision Tasks
Murat Bilgehan Ertan, Ronak Sahu, Phuong Ha Nguyen +2
We introduce ROAR (Robust Object Removal and Re-annotation), a scalable framework for privacy-preserving dataset obfuscation that eliminates sensitive objects instead of modifying…