Cockpit: A Practical Debugging Tool for the Training of Deep Neural Networks
arXiv:2102.06604
Abstract
When engineers train deep learning models, they are very much 'flying blind'. Commonly used methods for real-time training diagnostics, such as monitoring the train/test loss, are limited. Assessing a network's training process solely through these performance indicators is akin to debugging software without access to internal states through a debugger. To address this, we present Cockpit, a collection of instruments that enable a closer look into the inner workings of a learning machine, and a more informative and meaningful status report for practitioners. It facilitates the identification of learning phases and failure modes, like ill-chosen hyperparameters. These instruments leverage novel higher-order information about the gradient distribution and curvature, which has only recently become efficiently accessible. We believe that such a debugging tool, which we open-source for PyTorch, is a valuable help in troubleshooting the training process. By revealing new insights, it also more generally contributes to explainability and interpretability of deep nets.
(NeurIPS 2021) Main text: 13 pages, 6 figures, 1 table; Supplements: 23 pages, 13 figures, 1 table, 1 listing
References in corpus (13)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Gradient Descent Happens in a Tiny Subspace
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- Generalization in Deep Networks: The Role of Distance from Initialization
- The Early Phase of Neural Network Training
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
- A Study of Gradient Variance in Deep Learning
- Understanding Why Neural Networks Generalize Well Through GSNR of Parameters
- Making Coherence Out of Nothing At All: Measuring the Evolution of Gradient Alignment
- A Dynamic Sampling Adaptive-SGD Method for Machine Learning
- On regularization of gradient descent, layer imbalance and flat minima