5 citations · 8 across the 27 of their papers we have counts for
9 papers · 1 filter
CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs
Maryam Alshehyari, Dushyant Singh Chauhan, Samuele Poppi +3
Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchm…
Entropy-Gated Latent Recursion
Soham Bhattacharjee, Dushyant Singh Chauhan, Salem Lahlou +2
Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity from a single source: stochastic token-le…
A Gravitational Interpretation of Fine-Tuning Reversion
Samuele Poppi, Nils Lukas
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge,…
Collaborative Threshold Watermarking
Tameem Bakr, Anish Ambreth, Nils Lukas
In federated learning (FL), clients jointly train a model without sharing raw data. Because each participant invests data and compute, clients need mechanisms to later prove th…
Universal Backdoor Attacks
Benjamin Schneider, Nils Lukas, Florian Kerschbaum
Web-scraped datasets are vulnerable to data poisoning, which can be used for backdooring deep image classifiers during training. Since training on large datasets is expensive, a mo…
PTW: Pivotal Tuning Watermarking for Pre-Trained Image Generators
Nils Lukas, Florian Kerschbaum
Deepfakes refer to content synthesized using deep generators, which, when misused, have the potential to erode trust in digital media. Synthesizing high-quality deepfakes requires…