3 papers
cs.LG2025
Training on Plausible Counterfactuals Removes Spurious Correlations
Shpresim Sadiku, Kartikeya Chitranshi, Hiroshi Kera +1
Plausible counterfactual explanations (p-CFEs) are perturbations that minimally modify inputs to change classifier decisions while remaining plausible under the data distribution.…
cs.LG2025
S-CFE: Simple Counterfactual Explanations
Shpresim Sadiku, Moritz Wagner, Sai Ganesh Nagarajan +1
We study the problem of finding optimal sparse, manifold-aligned counterfactual explanations for classifiers. Canonically, this can be formulated as an optimization problem with mu…
cs.CV2025
GSE: Group-wise Sparse and Explainable Adversarial Attacks
Shpresim Sadiku, Moritz Wagner, Sebastian Pokutta
Sparse adversarial attacks fool deep neural networks (DNNs) through minimal pixel perturbations, often regularized by the norm. Recent efforts have replaced this norm with…