collaborators

6 papers

cs.AI2026

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

Wojciech Zarzecki, Jan Dubiński, Sebastian Cygert

Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for detecting training-data member…

cs.LG2026

Efficient LLM Moderation with Multi-Layer Latent Prototypes

Maciej ChrabÄ szcz, Filip Szatkowski, Bartosz Wójcik +3

Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time. Existing approaches suff…

cs.LG2026

Membership and Dataset Inference Attacks on Large Audio Generative Models

Jakub Proboszcz, Paweł Kochanski, Karol Korszun +5

Generative audio models, based on diffusion and autoregressive architectures, have advanced rapidly in both quality and expressiveness. This progress, however, raises pressing copy…

cs.CV2025

ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts

Patryk Będkowski, Jan Dubiński, Filip Szatkowski +3

Simulating detector responses is a crucial part of understanding the inner workings of particle collisions in the Large Hadron Collider at CERN. Such simulations are currently perf…

cs.LG2025

Radioactive Watermarks in Diffusion and Autoregressive Image Generative Models

Michel Meintz, Jan Dubiński, Franziska Boenisch +1

Image generative models have become increasingly popular, but training them requires large datasets that are costly to collect and curate. To circumvent these costs, some parties m…

cs.LG2025

CDI: Copyrighted Data Identification in Diffusion Models

Jan Dubiński, Antoni Kowalczuk, Franziska Boenisch +1

Diffusion Models (DMs) benefit from large and diverse datasets for their training. Since this data is often scraped from the Internet without permission from the data owners, this…