Adversarial flows: A gradient flow characterization of adversarial attacks
arXiv:2406.05376 · doi:10.1017/S0956792525100120
Abstract
A popular method to perform adversarial attacks on neuronal networks is the so-called fast gradient sign method and its iterative variant. In this paper, we interpret this method as an explicit Euler discretization of a differential inclusion, where we also show convergence of the discretization to the associated gradient flow. To do so, we consider the concept of p-curves of maximal slope in the case . We prove existence of -curves of maximum slope and derive an alternative characterization via differential inclusions. Furthermore, we also consider Wasserstein gradient flows for potential energies, where we show that curves in the Wasserstein space can be characterized by a representing measure on the space of curves in the underlying Banach space, which fulfill the differential inclusion. The application of our theory to the finite-dimensional setting is twofold: On the one hand, we show that a whole class of normalized gradient descent methods (in particular signed gradient descent) converge, up to subsequences, to the flow, when sending the step size to zero. On the other hand, in the distributional setting, we show that the inner optimization task of adversarial training objective can be characterized via -curves of maximum slope on an appropriate optimal transport space.
References in corpus (14)
- A consensus-based model for global optimization and its mean-field limit
- Training robust neural networks using Lipschitz bounds
- An easy proof of Jensen's theorem on the uniqueness of infinity harmonic functions
- Nonsmooth analysis of doubly nonlinear evolution equations
- Variational convergence of gradient flows and rate-independent evolutions in metric spaces
- Asymptotic Profiles of Nonlinear Homogeneous Evolution Equations of Gradient Flow Type
- Uniform Convergence Rates for Lipschitz Learning on Graphs
- Nonlinear Spectral Decompositions by Gradient Flows of One-Homogeneous Functionals
- The Geometry of Adversarial Training in Binary Classification
- CBX: Python and Julia packages for consensus-based interacting particle methods
- Gamma-convergence of a nonlocal perimeter arising in adversarial machine learning
- Ratio convergence rates for Euclidean first-passage percolation: Applications to the graph infinity Laplacian
- A mean curvature flow arising in adversarial training
- MirrorCBO: A consensus-based optimization method in the spirit of mirror descent