Linear Mode Connectivity and the Lottery Ticket Hypothesis
arXiv:1912.05671
Abstract
We study whether a neural network optimizes to the same, linearly connected minimum under different samples of SGD noise (e.g., random data order and augmentation). We find that standard vision models become stable to SGD noise in this way early in training. From then on, the outcome of optimization is determined to a linearly connected region. We use this technique to study iterative magnitude pruning (IMP), the procedure used by work on the lottery ticket hypothesis to identify subnetworks that could have trained in isolation to full accuracy. We find that these subnetworks only reach full accuracy when they are stable to SGD noise, which either occurs at initialization for small-scale settings (MNIST) or early in training for large-scale settings (ResNet-50 and Inception-v3 on ImageNet).
Published in ICML 2020. This submission subsumes arXiv:1903.01611 ("Stabilizing the Lottery Ticket Hypothesis" and "The Lottery Ticket Hypothesis at Scale")
References in corpus (14)
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- AMC: AutoML for Model Compression and Acceleration on Mobile Devices
- Rethinking the Value of Network Pruning
- Don't Decay the Learning Rate, Increase the Batch Size
- SNIP: Single-shot Network Pruning based on Connection Sensitivity
- The State of Sparsity in Deep Neural Networks
- Essentially No Barriers in Neural Network Energy Landscape
- Picking Winning Tickets Before Training by Preserving Gradient Flow
- Uniform convergence may be unable to explain generalization in deep learning
- Gradient Descent Happens in a Tiny Subspace
- Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP
- The Early Phase of Neural Network Training
- Winning the Lottery with Continuous Sparsification
- Topology and Geometry of Half-Rectified Network Optimization
Cited by in corpus (12)
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Pruning Convolutional Neural Networks with Self-Supervision
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Connecting Low-Loss Subspace for Personalized Federated Learning
- Riemannian Low-Rank Model Compression for Federated Learning with Over-the-Air Aggregation
- Pruning via Iterative Ranking of Sensitivity Statistics
- KinePose: A temporally optimized inverse kinematics technique for 6DOF human pose estimation with biomechanical constraints
- Train Flat, Then Compress: Sharpness-Aware Minimization Learns More Compressible Models
- On Lottery Tickets and Minimal Task Representations in Deep Reinforcement Learning
- Ultra-light deep MIR by trimming lottery tickets
- Weighted Ensemble Models Are Strong Continual Learners