Adversarial Parameter Defense by Multi-Step Risk Minimization
arXiv:2109.02889 · doi:10.1016/j.neunet.2021.08.022
Abstract
Previous studies demonstrate DNNs' vulnerability to adversarial examples and adversarial training can establish a defense to adversarial examples. In addition, recent studies show that deep neural networks also exhibit vulnerability to parameter corruptions. The vulnerability of model parameters is of crucial value to the study of model robustness and generalization. In this work, we introduce the concept of parameter corruption and propose to leverage the loss change indicators for measuring the flatness of the loss basin and the parameter robustness of neural network parameters. On such basis, we analyze parameter corruptions and propose the multi-step adversarial corruption algorithm. To enhance neural networks, we propose the adversarial parameter defense algorithm that minimizes the average risk of multiple adversarial parameter corruptions. Experimental results show that the proposed algorithm can improve both the parameter robustness and accuracy of neural networks.
Accepted to Neural Networks. A substantial journal extension of our previous conference paper arXiv:2006.05620
References in corpus (8)
- Explaining and Harnessing Adversarial Examples
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness
- Weight Poisoning Attacks on Pre-trained Models
- Regularizing Neural Networks via Adversarial Model Perturbation
- Improved Corruption Robust Algorithms for Episodic Reinforcement Learning