#adversarial training
6 papers · 1 filter
The Noise Premium in Adversarial Training for Kernel Regression
Yiling Xie, Xiaoming Huo
The paper analyzes adversarial training within reproducing kernel Hilbert spaces, deriving generalization bounds and showing how noise affects the trade‑off between robustness and…
FARI: Robust One-Step Inversion for Watermarking in Diffusion Models
Jindong Yang, Han Fang, Weiming Zhang +2
The paper introduces FARI, a fast one-step inversion method combined with lightweight adversarial LoRA fine-tuning to robustly extract watermarks from diffusion-generated images, a…
RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
Pushkal Kumar, Tucker Nielson, Tanish Kolhe +2
The paper introduces RAGuard, a two‑layer defense for retrieval‑augmented generation systems that combats factual corpus‑poisoning by adversarially fine‑tuning the retriever and ap…
GPT-Red: Automated Red Teaming via Self-Play at Scale
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15
The paper presents GPT-Red, an automated red‑teaming system that uses self‑play to generate novel prompt‑injection attacks against large language models and improve their robustnes…
Lights, Camera, Malfunction: When Illumination Robustness Leaves VLA Models Blind to Color
Marino Watanabe, Takami Sato, Kentaro Yoshioka
The paper studies how Vision‑Language‑Action (VLA) robot models fail under targeted spotlight illumination, reveals that common data augmentations cause them to ignore color, and i…
Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System
Ken Jon Miyachi, Dylan Uys
The paper introduces BitMind Forensics, a continuously updated deepfake detection system that uses an open adversarial competition to refresh its training data, and demonstrates st…