1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.LG2026★ 1 cited
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
Thomas Winninger, Boussad Addad, Katarzyna Kapusta
Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechani…
cs.CV2025
DiffGuard: Text-Based Safety Checker for Diffusion Models
Massine El Khader, Elias Al Bouzidi, Abdellah Oumida +6
Recent advances in Diffusion Models have enabled the generation of images from text, with powerful closed-source models like DALL-E and Midjourney leading the way. However, open-so…