The Lipschitz Constant of Self-Attention
arXiv:2006.04710
Abstract
Lipschitz constants of neural networks have been explored in various contexts in deep learning, such as provable adversarial robustness, estimating Wasserstein distance, stabilising training of GANs, and formulating invertible neural networks. Such works have focused on bounding the Lipschitz constant of fully connected or convolutional networks, composed of linear maps and pointwise non-linearities. In this paper, we investigate the Lipschitz constant of self-attention, a non-linear neural network module widely used in sequence modelling. We prove that the standard dot-product self-attention is not Lipschitz for unbounded input domain, and propose an alternative L2 self-attention that is Lipschitz. We derive an upper bound on the Lipschitz constant of L2 self-attention and provide empirical evidence for its asymptotic tightness. To demonstrate the practical relevance of our theoretical work, we formulate invertible self-attention and use it in a Transformer-based architecture for a character-level language modelling task.
References in corpus (3)
Cited by in corpus (14)
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation
- Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models
- Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
- DP-Forward: Fine-tuning and Inference on Language Models with Differential Privacy in Forward Pass
- A case for new neural network smoothness constraints
- A Mathematical Theory of Attention
- Lipschitz Normalization for Self-Attention Layers with Application to Graph Neural Networks
- Neural Spatio-Temporal Point Processes
- On the Regularity of Attention
- Probabilistic Transformers
- Implicit regularization of deep residual networks towards neural ODEs
- Coded-InvNet for Resilient Prediction Serving Systems
- Invertible Attention
- Use square root affinity to regress labels in semantic segmentation