Large-Margin Softmax Loss for Convolutional Neural Networks
arXiv:1612.02295
Abstract
Cross-entropy loss together with softmax is arguably one of the most common used supervision components in convolutional neural networks (CNNs). Despite its simplicity, popularity and excellent performance, the component does not explicitly encourage discriminative learning of features. In this paper, we propose a generalized large-margin softmax (L-Softmax) loss which explicitly encourages intra-class compactness and inter-class separability between learned features. Moreover, L-Softmax not only can adjust the desired margin but also can avoid overfitting. We also show that the L-Softmax loss can be optimized by typical stochastic gradient descent. Extensive experiments on four benchmark datasets demonstrate that the deeply-learned features with L-softmax loss become more discriminative, hence significantly boosting the performance on a variety of visual classification and verification tasks.
Published in ICML 2016 (with typo fixed)
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Improving neural networks by preventing co-adaptation of feature detectors
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Learning Face Representation from Scratch
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Deep Networks with Internal Selective Attention through Feedback Connections
Cited by in corpus (58)
- SphereFace: Deep Hypersphere Embedding for Face Recognition
- AdvHat: Real-world adversarial attack on ArcFace Face ID system
- Masked Face Recognition for Secure Authentication
- End2End Occluded Face Recognition by Masking Corrupted Features
- Rethinking Feature Discrimination and Polymerization for Large-scale Recognition
- Circle Loss: A Unified Perspective of Pair Similarity Optimization
- Parameter-Efficient Person Re-identification in the 3D Space
- von Mises-Fisher Mixture Model-based Deep learning: Application to Face Verification
- Remix: Rebalanced Mixup
- Towards duration robust weakly supervised sound event detection
- Minimum Margin Loss for Deep Face Recognition
- On the Geometry of Adversarial Examples
- Learning from Few Samples: A Survey
- Deep Gamblers: Learning to Abstain with Portfolio Theory
- Max-Mahalanobis Linear Discriminant Analysis Networks
- Negative Margin Matters: Understanding Margin in Few-shot Classification
- Iterative Learning with Open-set Noisy Labels
- Noisy Softmax: Improving the Generalization Ability of DCNN via Postponing the Early Softmax Saturation
- Joint Domain Alignment and Discriminative Feature Learning for Unsupervised Deep Domain Adaptation
- Robust Classification with Convolutional Prototype Learning
- Feature Transfer Learning for Deep Face Recognition with Under-Represented Data
- A Performance Comparison of Loss Functions for Deep Face Recognition
- iQIYI-VID: A Large Dataset for Multi-modal Person Identification
- Regular Polytope Networks
- Rethinking Feature Distribution for Loss Functions in Image Classification
- Regularizing Deep Networks with Semantic Data Augmentation
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Rapid detection and recognition of whole brain activity in a freely behaving Caenorhabditis elegans
- Learning to Anonymize Faces for Privacy Preserving Action Detection
- Large Margin Deep Networks for Classification
- Spectral Feature Transformation for Person Re-identification
- Viewpoint-Aware Loss with Angular Regularization for Person Re-Identification
- Multi-label Contrastive Predictive Coding
- Ring loss: Convex Feature Normalization for Face Recognition
- Ensemble Soft-Margin Softmax Loss for Image Classification
- Multi-shot Pedestrian Re-identification via Sequential Decision Making
- Scalable Angular Discriminative Deep Metric Learning for Face Recognition
- Accelerating Large Scale Knowledge Distillation via Dynamic Importance Sampling
- Reconstruction of Simulation-Based Physical Field by Reconstruction Neural Network Method
- More Information Supervised Probabilistic Deep Face Embedding Learning
- Focal Inferential Infusion Coupled with Tractable Density Discrimination for Implicit Hate Detection
- Adaptive Prototypical Networks with Label Words and Joint Representation Learning for Few-Shot Relation Classification
- Normalization Before Shaking Toward Learning Symmetrically Distributed Representation Without Margin in Speech Emotion Recognition
- An Attention Model for group-level emotion recognition
- Distributed Map Classification using Local Observations
- CAMRI Loss: Improving Recall of a Specific Class without Sacrificing Accuracy
- Adaptive Discriminative Regularization for Visual Classification
- Adma: A Flexible Loss Function for Neural Networks
- Tackling Early Sparse Gradients in Softmax Activation Using Leaky Squared Euclidean Distance
- Anchor-based Nearest Class Mean Loss for Convolutional Neural Networks
- A Deep Learning Framework using Passive WiFi Sensing for Respiration Monitoring
- Generalized Categorisation of Digital Pathology Whole Image Slides using Unsupervised Learning
- On Symmetry and Initialization for Neural Networks
- Push for Center Learning via Orthogonalization and Subspace Masking for Person Re-Identification
- GIM: Gaussian Isolation Machines
- The Best of Both Worlds: a Framework for Combining Degradation Prediction with High Performance Super-Resolution Networks
- Contrastive Forward-Forward: A Training Algorithm of Vision Transformer
- Exponential Discriminative Metric Embedding in Deep Learning