Publications (13)
Detecting Benchmark Contamination Through Watermarking
Tom Sander, Pierre Fernandez, Saeed Mahloujifar +2
Benchmark contamination poses a significant challenge to the reliability of Large Language Models (LLMs) evaluations, as it is difficult to assert whether a model has been trained…
How Good is Post-Hoc Watermarking With Language Model Rephrasing?
Pierre Fernandez, Tom Sander, Hady Elsahar +6
Generation-time text watermarking embeds statistical signals into text for traceability of AI-generated content. We explore *post-hoc watermarking* where an LLM rewrites existing t…
Pixel Seal: Adversarial-only training for invisible image and video watermarking
Tomáš SouÄek, Pierre Fernandez, Hady Elsahar +5
Invisible watermarking is essential for tracing the provenance of digital content. However, training state-of-the-art models remains notoriously difficult, with current approaches…
Watermark Anything with Localized Messages
Tom Sander, Pierre Fernandez, Alain Durmus +2
Image watermarking methods are not tailored to handle small watermarked areas. This restricts applications in real-world scenarios where parts of the image may come from different…
Learning to Watermark in the Latent Space of Generative Models
Sylvestre-Alvise Rebuffi, Tuan Tran, Valeriu Lacatusu +6
Existing approaches for watermarking AI-generated images often rely on post-hoc methods applied in pixel space, introducing computational overhead and potential visual artifacts. I…
Implicit Bias in Noisy-SGD: With Applications to Differentially Private Training
Tom Sander, Maxime Sylvestre, Alain Durmus
Training Deep Neural Networks (DNNs) with small batches using Stochastic Gradient Descent (SGD) yields superior test performance compared to larger batches. The specific noise stru…
Watermarking Makes Language Models Radioactive
Tom Sander, Pierre Fernandez, Alain Durmus +2
We investigate the radioactivity of text generated by large language models (LLM), i.e. whether it is possible to detect that such synthetic input was used to train a subsequent LL…
Differentially Private Representation Learning via Image Captioning
Tom Sander, Yaodong Yu, Maziar Sanjabi +4
Differentially private (DP) machine learning is considered the gold-standard solution for training a model from sensitive data while still preserving privacy. However, a major barr…
TAN Without a Burn: Scaling Laws of DP-SGD
Tom Sander, Pierre Stock, Alexandre Sablayrolles
Differentially Private methods for training Deep Neural Networks (DNNs) have progressed recently, in particular with the use of massive batches and aggregated data augmentations fo…
The Moon's Many Faces: A Single Unified Transformer for Multimodal Lunar Reconstruction
Tom Sander, Moritz Tenthoff, Kay Wohlfarth +1
Multimodal learning is an emerging research topic across multiple disciplines but has rarely been applied to planetary science. In this contribution, we propose a single, unified t…
TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
Tom Sander, Hongyan Chang, Tomáš SouÄek +10
We introduce TextSeal, a state-of-the-art watermark for large language models. Building on Gumbel-max sampling, TextSeal introduces dual-key generation to restore output diversity,…
Log-normal Mutations and their Use in Detecting Surreptitious Fake Images
Ismail Labiad, Thomas Bäck, Pierre Fernandez +5
In many cases, adversarial attacks are based on specialized algorithms specifically dedicated to attacking automatic image classifiers. These algorithms perform well, thanks to an…
Verifiably grounded machine interpretation of lunar geology
Tom Sander, Kay Wohlfarth, Christian Wöhler
Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we present a step toward an automated "machine intelligen…