Loss Patterns of Neural Networks
arXiv:1910.03867
Abstract
We present multi-point optimization: an optimization technique that allows to train several models simultaneously without the need to keep the parameters of each one individually. The proposed method is used for a thorough empirical analysis of the loss landscape of neural networks. By extensive experiments on FashionMNIST and CIFAR10 datasets we demonstrate two things: 1) loss surface is surprisingly diverse and intricate in terms of landscape patterns it contains, and 2) adding batch normalization makes it more smooth. Source code to reproduce all the reported results is available on GitHub: https://github.com/universome/loss-patterns.
References in corpus (4)
Cited by in corpus (8)
- Towards a Mathematical Understanding of Neural Network-Based Machine Learning: what we know and what we don't
- On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them
- Loss Surface Simplexes for Mode Connecting Volumes and Fast Ensembling
- Learn Faster and Forget Slower via Fast and Stable Task Adaptation
- Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
- LossPlot: A Better Way to Visualize Loss Landscapes
- Embedding Principle: a hierarchical structure of loss landscape of deep neural networks
- Deep Learning in Target Space