On The Distribution of Penultimate Activations of Classification Networks
arXiv:2107.01900
Abstract
This paper studies probability distributions of penultimate activations of classification networks. We show that, when a classification network is trained with the cross-entropy loss, its final classification layer forms a Generative-Discriminative pair with a generative classifier based on a specific distribution of penultimate activations. More importantly, the distribution is parameterized by the weights of the final fully-connected layer, and can be considered as a generative model that synthesizes the penultimate activations without feeding input data. We empirically demonstrate that this generative model enables stable knowledge distillation in the presence of domain shift, and can transfer knowledge from a classifier to variational autoencoders and generative adversarial networks for class-conditional image generation.
8 pages, UAI 2021, The first two authors equally contributed
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- FitNets: Hints for Thin Deep Nets
- Semi-Supervised Learning with Deep Generative Models
- Understanding Neural Networks Through Deep Visualization
- von Mises-Fisher Mixture Model-based Deep learning: Application to Face Verification
- Radial and Directional Posteriors for Bayesian Neural Networks