Knowledge Transfer with Jacobian Matching
arXiv:1803.00443
Abstract
Classical distillation methods transfer representations from a "teacher" neural network to a "student" network by matching their output activations. Recent methods also match the Jacobians, or the gradient of output activations with the input. However, this involves making some ad hoc decisions, in particular, the choice of the loss function. In this paper, we first establish an equivalence between Jacobian matching and distillation with input noise, from which we derive appropriate loss functions for Jacobian matching. We then rely on this analysis to apply Jacobian matching to transfer learning by establishing equivalence of a recent transfer learning procedure to distillation. We then show experimentally on standard image datasets that Jacobian-based penalties improve distillation, robustness to noisy inputs, and transfer learning.
Cited by in corpus (26)
- Self-Distillation Amplifies Regularization in Hilbert Space
- Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons
- Efficient Convolutional Neural Networks for Depth-Based Multi-Person Pose Estimation
- Self-Distillation as Instance-Specific Label Smoothing
- Regularizing Class-wise Predictions via Self-knowledge Distillation
- ShrinkTeaNet: Million-scale Lightweight Face Recognition via Shrinking Teacher-Student Networks
- Domain Adaptation without Source Data
- Diversity Matters When Learning From Ensembles
- Few Sample Knowledge Distillation for Efficient Network Compression
- A Comprehensive Overhaul of Feature Distillation
- Gradients as Features for Deep Representation Learning
- Oblique Decision Trees from Derivatives of ReLU Networks
- Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior Knowledge
- Network Transplanting
- Even your Teacher Needs Guidance: Ground-Truth Targets Dampen Regularization Imposed by Self-Distillation
- Multi-step Estimation for Gradient-based Meta-learning
- Knowledge Distillation Meets Self-Supervision
- Matching Guided Distillation
- Neural Parameter Allocation Search
- Prune Your Model Before Distill It
- Similarity Transfer for Knowledge Distillation
- Network Transplanting (extended abstract)
- Collaborative Group Learning
- A Survey on Green Deep Learning
- Student-Teacher Learning from Clean Inputs to Noisy Inputs
- Locally Linear Region Knowledge Distillation