Prototype Memory for Large-scale Face Representation Learning
arXiv:2105.02103 · doi:10.1109/ACCESS.2022.3146059
Abstract
Face representation learning using datasets with a massive number of identities requires appropriate training methods. Softmax-based approach, currently the state-of-the-art in face recognition, in its usual "full softmax" form is not suitable for datasets with millions of persons. Several methods, based on the "sampled softmax" approach, were proposed to remove this limitation. These methods, however, have a set of disadvantages. One of them is a problem of "prototype obsolescence": classifier weights (prototypes) of the rarely sampled classes receive too scarce gradients and become outdated and detached from the current encoder state, resulting in incorrect training signals. This problem is especially serious in ultra-large-scale datasets. In this paper, we propose a novel face representation learning model called Prototype Memory, which alleviates this problem and allows training on a dataset of any size. Prototype Memory consists of the limited-size memory module for storing recent class prototypes and employs a set of algorithms to update it in appropriate way. New class prototypes are generated on the fly using exemplar embeddings in the current mini-batch. These prototypes are enqueued to the memory and used in a role of classifier weights for softmax classification-based training. To prevent obsolescence and keep the memory in close connection with the encoder, prototypes are regularly refreshed, and oldest ones are dequeued and disposed of. Prototype Memory is computationally efficient and independent of dataset size. It can be used with various loss functions, hard example mining algorithms and encoder architectures. We prove the effectiveness of the proposed model by extensive experiments on popular face recognition benchmarks.
References in corpus (23)
- Distilling the Knowledge in a Neural Network
- Transformers in Vision: A Survey
- RetinaFace: Single-stage Dense Face Localisation in the Wild
- Cross-Age LFW: A Database for Studying Cross-Age Face Recognition in Unconstrained Environments
- SFace: Sigmoid-Constrained Hypersphere Loss for Robust Face Recognition
- Rethinking Feature Discrimination and Polymerization for Large-scale Recognition
- Partial FC: Training 10 Million Identities on a Single Machine
- Face Recognition via Centralized Coordinate Learning
- Loss Function Search for Face Recognition
- ShrinkTeaNet: Million-scale Lightweight Face Recognition via Shrinking Teacher-Student Networks
- Cross-Resolution Face Recognition via Prior-Aided Face Hallucination and Residual Knowledge Distillation
- The Elements of End-to-end Deep Face Recognition: A Survey of Recent Advances
- Attribute Adaptive Margin Softmax Loss using Privileged Information
- Extreme Classification via Adversarial Softmax Approximation
- Masked Face Recognition: Human vs. Machine
- NPT-Loss: A Metric Loss with Implicit Mining for Face Recognition
- Fast and Reliable Probabilistic Face Embeddings in the Wild
- TAPAS: Two-pass Approximate Adaptive Sampling for Softmax
- MarginDistillation: distillation for margin-based softmax
- MultiFace: A Generic Training Mechanism for Boosting Face Recognition Performance
- Self-organized Hierarchical Softmax
- Inter-class Discrepancy Alignment for Face Recognition
- More Information Supervised Probabilistic Deep Face Embedding Learning