A Regularized Framework for Sparse and Structured Neural Attention
arXiv:1705.07704
Abstract
Modern neural networks are often augmented with an attention mechanism, which tells the network where to focus within the input. We propose in this paper a new framework for sparse and structured attention, building upon a smoothed max operator. We show that the gradient of this operator defines a mapping from real values to probabilities, suitable as an attention mechanism. Our framework includes softmax and a slight generalization of the recently-proposed sparsemax as special cases. However, we also show how our framework can incorporate modern structured penalties, resulting in more interpretable attention mechanisms, that focus on entire segments or groups of an input. We derive efficient algorithms to compute the forward and backward passes of our attention mechanisms, enabling their use in a neural network trained with backpropagation. To showcase their potential as a drop-in replacement for existing ones, we evaluate our attention mechanisms on three large-scale tasks: textual entailment, machine translation, and sentence summarization. Our attention mechanisms improve interpretability without sacrificing performance; notably, on textual entailment and summarization, we outperform the standard attention mechanisms based on softmax and sparsemax.
In proceedings of NeurIPS 2017; added errata
Cited by in corpus (22)
- SparseMAP: Differentiable Sparse Structured Inference
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular Data
- Efficient and Modular Implicit Differentiation
- Implicit differentiation of Lasso-type models for hyperparameter optimization
- Adaptively Sparse Transformers
- Fast On-the-fly Retraining-free Sparsification of Convolutional Neural Networks
- Trilevel Neural Architecture Search for Efficient Single Image Super-Resolution
- Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
- Sparse Sequence-to-Sequence Models
- On Sparsifying Encoder Outputs in Sequence-to-Sequence Models
- Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS
- Improving Textual Network Embedding with Global Attention via Optimal Transport
- Interpretable Structured Learning with Sparse Gated Sequence Encoder for Protein-Protein Interaction Prediction
- Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation
- Learning Classifiers with Fenchel-Young Losses: Generalized Entropies, Margins, and Algorithms
- Sparse Attention with Linear Units
- Not All Attention Is Needed: Gated Attention Network for Sequence Data
- Scaling Up Multiagent Reinforcement Learning for Robotic Systems: Learn an Adaptive Sparse Communication Graph
- Yet Another Representation of Binary Decision Trees: A Mathematical Demonstration
- A Survey on Green Deep Learning
- Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop
- Decision Machines: Congruent Decision Trees