Visualizing the Loss Landscape of Neural Nets
arXiv:1712.09913
Abstract
Neural network training relies on our ability to find "good" minimizers of highly non-convex loss functions. It is well-known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, optimizer) produce minimizers that generalize better. However, the reasons for these differences, and their effects on the underlying loss landscape, are not well understood. In this paper, we explore the structure of neural loss functions, and the effect of loss landscapes on generalization, using a range of visualization methods. First, we introduce a simple "filter normalization" method that helps us visualize loss function curvature and make meaningful side-by-side comparisons between loss functions. Then, using a variety of visualizations, we explore how network architecture affects the loss landscape, and how training parameters affect the shape of minimizers.
NIPS 2018 (extended version, 10.5 pages), code is available at https://github.com/tomgoldstein/loss-landscape
References in corpus (11)
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Convergence Analysis of Two-layer Neural Networks with ReLU Activation
- Qualitatively characterizing neural network optimization problems
- The loss surface of deep and wide neural networks
- Diverse Neural Network Learns True Target Functions
- Exponentially vanishing sub-optimal local minima in multilayer neural networks
- Global optimality conditions for deep neural networks
- Theory II: Landscape of the Empirical Risk in Deep Learning
- Local minima in training of neural networks
- Exploring loss function topology with cyclical learning rates
Cited by in corpus (93)
- TransMorph: Transformer for unsupervised medical image registration
- FedALA: Adaptive Local Aggregation for Personalized Federated Learning
- A Primer on Motion Capture with Deep Learning: Principles, Pitfalls and Perspectives
- Machine-Learning-Based Diagnostics of EEG Pathology
- A Machine Learning Framework for Solving High-Dimensional Mean Field Game and Mean Field Control Problems
- Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
- Essentially No Barriers in Neural Network Energy Landscape
- How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
- Review: Deep Learning in Electron Microscopy
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- Machine-Learning-Based Multiple Abnormality Prediction with Large-Scale Chest Computed Tomography Volumes
- End-to-end learning for music audio tagging at scale
- The Global Landscape of Neural Networks: An Overview
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
- A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
- Promises and pitfalls of deep neural networks in neuroimaging-based psychiatric research
- Nostalgic Adam: Weighting more of the past gradients when designing the adaptive learning rate
- GraphAIR: Graph Representation Learning with Neighborhood Aggregation and Interaction
- Generative model benchmarks for superconducting qubits
- End-to-end optimization of coherent optical communications over the split-step Fourier method guided by the nonlinear Fourier transform theory
- Learning macroscopic internal variables and history dependence from microscopic models
- Development of Skip Connection in Deep Neural Networks for Computer Vision and Medical Image Analysis: A Survey
- Auto-Ensemble: An Adaptive Learning Rate Scheduling based Deep Learning Model Ensembling
- U(1) symmetric recurrent neural networks for quantum state reconstruction
- Knowledge Distillation via Route Constrained Optimization
- Investigating and Mitigating Failure Modes in Physics-informed Neural Networks (PINNs)
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- Laplacian Smoothing Gradient Descent
- Mathematical Models of Overparameterized Neural Networks
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- Laplace HypoPINN: Physics-Informed Neural Network for hypocenter localization and its predictive uncertainty
- Privacy-Preserving Ensemble Infused Enhanced Deep Neural Network Framework for Edge Cloud Convergence
- Data efficiency and extrapolation trends in neural network interatomic potentials
- A Scale Invariant Flatness Measure for Deep Network Minima
- deep-significance - Easy and Meaningful Statistical Significance Testing in the Age of Neural Networks
- Near-optimal control of dynamical systems with neural ordinary differential equations
- SAT: Improving Adversarial Training via Curriculum-Based Loss Smoothing
- Towards Assessing the Synthetic-to-Measured Adversarial Vulnerability of SAR ATR
- Multi-class Classification without Multi-class Labels
- Deep learning via message passing algorithms based on belief propagation
- Quantitative Propagation of Chaos for SGD in Wide Neural Networks
- An Unsupervised Framework for Dynamic Health Indicator Construction and Its Application in Rolling Bearing Prognostics
- Towards Understanding Generalization in Gradient-Based Meta-Learning
- Building Surrogate Models of Nuclear Density Functional Theory with Gaussian Processesand Autoencoders
- LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors
- NAC-TCN: Temporal Convolutional Networks with Causal Dilated Neighborhood Attention for Emotion Understanding
- Training models using forces computed by stochastic electronic structure methods
- Visualizing high-dimensional loss landscapes with Hessian directions
- Precision Highway for Ultra Low-Precision Quantization
- Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
- Visualized Insights into the Optimization Landscape of Fully Convolutional Networks
- SIRe-Networks: Convolutional Neural Networks Architectural Extension for Information Preservation via Skip/Residual Connections and Interlaced Auto-Encoders
- Transfer Learning Enhanced Full Waveform Inversion
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- DHA: End-to-End Joint Optimization of Data Augmentation Policy, Hyper-parameter and Architecture
- Field-level simulation-based inference with galaxy catalogs: the impact of systematic effects
- Equivariant Neural Networks for Spin Dynamics Simulations of Itinerant Magnets
- Learning ReLU Networks via Alternating Minimization
- Understanding and Accelerating Neural Architecture Search with Training-Free and Theory-Grounded Metrics
- Predicting Change, Not States: An Alternate Framework for Neural PDE Surrogates
- Perturbated Gradients Updating within Unit Space for Deep Learning
- Enhancing Transformers without Self-supervised Learning: A Loss Landscape Perspective in Sequential Recommendation
- On generalization bounds for deep networks based on loss surface implicit regularization
- On Visual Hallmarks of Robustness to Adversarial Malware
- Evolutionary Augmentation Policy Optimization for Self-supervised Learning
- Enhancing Variational Quantum Circuit Training: An Improved Neural Network Approach for Barren Plateau Mitigation
- Understanding the Functional Roles of Modelling Components in Spiking Neural Networks
- Universal characteristics of deep neural network loss surfaces from random matrix theory
- Deep Curvature Suite
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy Compression
- ConvNets for Counting: Object Detection of Transient Phenomena in Steelpan Drums
- FedSC: Federated Learning with Semantic-Aware Collaboration
- MEAT: Median-Ensemble Adversarial Training for Improving Robustness and Generalization
- On the Evolution of Neuron Communities in a Deep Learning Architecture
- Pro-KD: Progressive Distillation by Following the Footsteps of the Teacher
- qLEET: Visualizing Loss Landscapes, Expressibility, Entangling Power and Training Trajectories for Parameterized Quantum Circuits
- SuperNet -- An efficient method of neural networks ensembling
- Constrained Linear Data-feature Mapping for Image Classification
- Extrapolation for Large-batch Training in Deep Learning
- ORQVIZ: Visualizing High-Dimensional Landscapes in Variational Quantum Algorithms
- Privately Learning Subspaces
- Traversing the noise of dynamic mini-batch sub-sampled loss functions: A visual guide
- Universal Performance Gap of Neural Quantum States Applied to the Hofstadter-Bose-Hubbard Model
- Enhance Diffusion to Improve Robust Generalization
- Gradient-only line searches to automatically determine learning rates for a variety of stochastic training algorithms
- HERO: Hessian-Enhanced Robust Optimization for Unifying and Improving Generalization and Quantization Performance
- Ensemble Feature for Person Re-Identification
- An Information-theoretic Visual Analysis Framework for Convolutional Neural Networks
- Analysis of Atomistic Representations Using Weighted Skip-Connections
- Riemannian Laplace approximations for Bayesian neural networks
- Data optimization for large batch distributed training of deep neural networks
- Left ventricle segmentation By modelling uncertainty in prediction of deep convolutional neural networks and adaptive thresholding inference
- TOFU: Towards Obfuscated Federated Updates by Encoding Weight Updates into Gradients from Proxy Data