From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification
arXiv:1602.02068
Abstract
We propose sparsemax, a new activation function similar to the traditional softmax, but able to output sparse probabilities. After deriving its properties, we show how its Jacobian can be efficiently computed, enabling its use in a network trained with backpropagation. Then, we propose a new smooth and convex loss function which is the sparsemax analogue of the logistic loss. We reveal an unexpected connection between this new loss and the Huber classification loss. We obtain promising empirical results in multi-label classification problems and in attention-based neural networks for natural language inference. For the latter, we achieve a similar performance as the traditional softmax, but with a selective, more compact, attention focus.
Minor corrections
References in corpus (3)
Cited by in corpus (75)
- Deep Neural Networks and Tabular Data: A Survey
- Attention in Natural Language Processing
- From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification
- On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning
- Decision-Focused Learning: Foundations, State of the Art, Benchmark and Future Opportunities
- OpenNMT: Neural Machine Translation Toolkit
- Differentiable Dynamic Programming for Structured Prediction and Attention
- SparseMAP: Differentiable Sparse Structured Inference
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular Data
- Contingency-Aware Exploration in Reinforcement Learning
- ConsRec: Learning Consensus Behind Interactions for Group Recommendation
- NeuralTailor: Reconstructing Sewing Pattern Structures from 3D Point Clouds of Garments
- Efficient and Modular Implicit Differentiation
- A Tutorial on Deep Latent Variable Models of Natural Language
- An Interpretable Knowledge Transfer Model for Knowledge Base Completion
- A Random Block-Coordinate Douglas-Rachford Splitting Method with Low Computational Complexity for Binary Logistic Regression
- On Controllable Sparse Alternatives to Softmax
- Gradient Estimation with Stochastic Softmax Tricks
- From English To Foreign Languages: Transferring Pre-trained Language Models
- A multi-label, dual-output deep neural network for automated bug triaging
- Sigsoftmax: Reanalysis of the Softmax Bottleneck
- Adaptively Sparse Transformers
- A Survey on Document-level Neural Machine Translation: Methods and Evaluation
- Sparse and Continuous Attention Mechanisms
- DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
- ProtoAttend: Attention-Based Prototypical Learning
- Trilevel Neural Architecture Search for Efficient Single Image Super-Resolution
- Dual-interest Factorization-heads Attention for Sequential Recommendation
- Pay Attention to the cough: Early Diagnosis of COVID-19 using Interpretable Symptoms Embeddings with Cough Sound Signal Processing
- The Limited Multi-Label Projection Layer
- Fast Differentiable Sorting and Ranking
- Towards Interpretable Sparse Graph Representation Learning with Laplacian Pooling
- Differentially Private Query Release Through Adaptive Projection
- Aspect-augmented Adversarial Networks for Domain Adaptation
- Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
- Optimizing for Interpretability in Deep Neural Networks with Tree Regularization
- Regional Tree Regularization for Interpretability in Black Box Models
- A Measure-Theoretic Characterization of Tight Language Models
- SSN: Learning Sparse Switchable Normalization via SparsestMax
- LP-SparseMAP: Differentiable Relaxed Optimization for Sparse Structured Prediction
- Learning with Fenchel-Young Losses
- Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS
- Towards Decoding as Continuous Optimization in Neural Machine Translation
- FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
- Ensemble Soft-Margin Softmax Loss for Image Classification
- Sparse Text Generation
- Interpretable Structured Learning with Sparse Gated Sequence Encoder for Protein-Protein Interaction Prediction
- Know Your Limits: Uncertainty Estimation with ReLU Classifiers Fails at Reliable OOD Detection
- Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders
- Learning Classifiers with Fenchel-Young Losses: Generalized Entropies, Margins, and Algorithms
- A Comparative Study of Deep Learning Loss Functions for Multi-Label Remote Sensing Image Classification
- FastAdaBelief: Improving Convergence Rate for Belief-based Adaptive Optimizers by Exploiting Strong Convexity
- Who2com: Collaborative Perception via Learnable Handshake Communication
- Why Attentions May Not Be Interpretable?
- Attention augmented differentiable forest for tabular data
- Transfer Learning with Pre-trained Conditional Generative Models
- Sparse Attention with Linear Units
- Scaling sparsemax based channel selection for speech recognition with ad-hoc microphone arrays
- Graphmax for Text Generation
- Context-aware Non-linear and Neural Attentive Knowledge-based Models for Grade Prediction
- Yet Another Representation of Binary Decision Trees: A Mathematical Demonstration
- Sparse Communication via Mixed Distributions
- Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement Learning
- Convergent Graph Solvers
- Sparse Continuous Distributions and Fenchel-Young Losses
- Learning Sparsity of Representations with Discrete Latent Variables
- Kernel Deformed Exponential Families for Sparse Continuous Attention
- Deep Unsupervised Drum Transcription
- Optimal Approximation -- Smoothness Tradeoffs for Soft-Max Functions
- Supervised Tree-Wasserstein Distance
- Learning sparse transformations through backpropagation
- Decision Machines: Congruent Decision Trees
- ParaQG: A System for Generating Questions and Answers from Paragraphs
- Pretext Tasks selection for multitask self-supervised speech representation learning
- A Survey on Green Deep Learning