Think Global, Act Local: Relating DNN generalisation and node-level SNR
arXiv:2002.04687
Abstract
The reasons behind good DNN generalisation remain an open question. In this paper we explore the problem by looking at the Signal-to-Noise Ratio of nodes in the network. Starting from information theory principles, it is possible to derive an expression for the SNR of a DNN node output. Using this expression we construct figures-of-merit that quantify how well the weights of a node optimise SNR (or, equivalently, information rate). Applying these figures-of-merit, we give examples indicating that weight sets that promote good SNR performance also exhibit good generalisation. In addition, we are able to identify the qualities of weight sets that exhibit good SNR behaviour and hence promote good generalisation. This leads to a discussion of how these results relate to network training and regularisation. Finally, we identify some ways that these observations can be used in training design.
15 pages, 5 figures; for associated colab files see http://github.com/pnorridge/think-global-act-local/settings
References in corpus (5)
- Improving neural networks by preventing co-adaptation of feature detectors
- Understanding deep learning requires rethinking generalization
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Stochastic Mirror Descent on Overparameterized Nonlinear Models: Convergence, Implicit Regularization, and Generalization
- Information Bottleneck and its Applications in Deep Learning