S2SD: Simultaneous Similarity-based Self-Distillation for Deep Metric Learning
arXiv:2009.08348
Abstract
Deep Metric Learning (DML) provides a crucial tool for visual similarity and zero-shot applications by learning generalizing embedding spaces, although recent work in DML has shown strong performance saturation across training objectives. However, generalization capacity is known to scale with the embedding space dimensionality. Unfortunately, high dimensional embeddings also create higher retrieval cost for downstream applications. To remedy this, we propose \emph{Simultaneous Similarity-based Self-distillation (S2SD). S2SD extends DML with knowledge distillation from auxiliary, high-dimensional embedding and feature spaces to leverage complementary context during training while retaining test-time cost and with negligible changes to the training time. Experiments and ablations across different objectives and standard benchmarks show S2SD offers notable improvements of up to 7% in Recall@1, while also setting a new state-of-the-art. Code available at https://github.com/MLforHealth/S2SD.
Accepted to ICML2021
References in corpus (11)
- Distilling the Knowledge in a Neural Network
- A Simple Framework for Contrastive Learning of Visual Representations
- Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
- Multi-Level Variational Autoencoder: Learning Disentangled Representations from Grouped Observations
- Rethinking Few-Shot Image Classification: a Good Embedding Is All You Need?
- Revisiting Training Strategies and Generalization Performance in Deep Metric Learning
- Self-supervised Knowledge Distillation for Few-shot Learning
- Sharing Matters for Generalization in Deep Metric Learning
- Transferring Inductive Biases through Knowledge Distillation
- MetaDistiller: Network Self-Boosting via Meta-Learned Top-Down Distillation