The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
arXiv:1803.03635
Abstract
Neural network pruning techniques can reduce the parameter counts of trained networks by over 90%, decreasing storage requirements and improving computational performance of inference without compromising accuracy. However, contemporary experience is that the sparse architectures produced by pruning are difficult to train from the start, which would similarly improve training performance. We find that a standard pruning technique naturally uncovers subnetworks whose initializations made them capable of training effectively. Based on these results, we articulate the "lottery ticket hypothesis:" dense, randomly-initialized, feed-forward networks contain subnetworks ("winning tickets") that - when trained in isolation - reach test accuracy comparable to the original network in a similar number of iterations. The winning tickets we find have won the initialization lottery: their connections have initial weights that make training particularly effective. We present an algorithm to identify winning tickets and a series of experiments that support the lottery ticket hypothesis and the importance of these fortuitous initializations. We consistently find winning tickets that are less than 10-20% of the size of several fully-connected and convolutional feed-forward architectures for MNIST and CIFAR10. Above this size, the winning tickets that we find learn faster than the original network and reach higher test accuracy.
ICLR camera ready
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Improving neural networks by preventing co-adaptation of feature detectors
- Understanding deep learning requires rethinking generalization
- Rethinking the Value of Network Pruning
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Pruning Filters for Efficient ConvNets
- A Closer Look at Memorization in Deep Networks
- Deep Rewiring: Training very sparse deep networks
- Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach
- Training Skinny Deep Neural Networks with Iterative Hard Thresholding Methods
Cited by in corpus (269)
- Convolutional Neural Networks as a Model of the Visual System: Past, Present, and Future
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Machine Learning and Deep Learning -- A review for Ecologists
- Multi-Task Learning with Deep Neural Networks: A Survey
- A deep-learning-based surrogate model for data assimilation in dynamic subsurface flow problems
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Learning Representations for Neural Network-Based Classification Using the Information Bottleneck Principle
- HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous Clients
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Plan-Structured Deep Neural Network Models for Query Performance Prediction
- Optimization for deep learning: theory and algorithms
- The Modern Mathematics of Deep Learning
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- Machine learning with neural networks
- Interpreting Deep Learning-Based Networking Systems
- Unlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding Emotion
- Deep Polynomial Neural Networks
- Logic Explained Networks
- Learning Sparse Networks Using Targeted Dropout
- LotteryFL: Personalized and Communication-Efficient Federated Learning with Lottery Ticket Hypothesis on Non-IID Datasets
- Proving the Lottery Ticket Hypothesis: Pruning is All You Need
- Filter Pruning by Switching to Neighboring CNNs with Good Attributes
- Convergence of Artificial Intelligence and High Performance Computing on NSF-supported Cyberinfrastructure
- Neuron Shapley: Discovering the Responsible Neurons
- How Can We Be So Dense? The Benefits of Using Highly Sparse Representations
- Advanced Dropout: A Model-free Methodology for Bayesian Dropout Optimization
- Great Power, Great Responsibility: Recommendations for Reducing Energy for Training Language Models
- ConfusionFlow: A model-agnostic visualization for temporal analysis of classifier confusion
- Reducing Computational Complexity of Neural Networks in Optical Channel Equalization: From Concepts to Implementation
- Recent Advances of Differential Privacy in Centralized Deep Learning: A Systematic Survey
- MEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks
- QKD: Quantization-aware Knowledge Distillation
- Superposition of many models into one
- Model Pruning Enables Localized and Efficient Federated Learning for Yield Forecasting and Data Sharing
- Only Train Once: A One-Shot Neural Network Training And Pruning Framework
- What Kinds of Functions do Deep Neural Networks Learn? Insights from Variational Spline Theory
- Knowledge Distillation approach towards Melanoma Detection
- Complexity-Driven CNN Compression for Resource-constrained Edge AI
- On-device Training: A First Overview on Existing Systems
- Control the model sign problem via path optimization method: Monte-Carlo approach to QCD effective model with Polyakov loop
- OptEmbed: Learning Optimal Embedding Table for Click-through Rate Prediction
- Mixed Dimension Embeddings with Application to Memory-Efficient Recommendation Systems
- LOss-Based SensiTivity rEgulaRization: towards deep sparse neural networks
- Training independent subnetworks for robust prediction
- Spiking neural networks trained with backpropagation for low power neuromorphic implementation of voice activity detection
- Conservative Sparse Neural Network Embedded Frequency-Constrained Unit Commitment With Distributed Energy Resources
- Policy Manifold Search: Exploring the Manifold Hypothesis for Diversity-based Neuroevolution
- Understanding Neural Networks and Individual Neuron Importance via Information-Ordered Cumulative Ablation
- Rethinking Weight Decay For Efficient Neural Network Pruning
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- CNN Filter DB: An Empirical Investigation of Trained Convolutional Filters
- Optimal Lottery Tickets via SubsetSum: Logarithmic Over-Parameterization is Sufficient
- Geometric framework to predict structure from function in neural networks
- How fine can fine-tuning be? Learning efficient language models
- DP-BART for Privatized Text Rewriting under Local Differential Privacy
- Task dependent Deep LDA pruning of neural networks
- Knowledge Distillation Methods for Efficient Unsupervised Adaptation Across Multiple Domains
- Born-Again Tree Ensembles
- End-to-End Supermask Pruning: Learning to Prune Image Captioning Models
- DiPair: Fast and Accurate Distillation for Trillion-Scale Text Matching and Pair Modeling
- Loss-Aware Automatic Selection of Structured Pruning Criteria for Deep Neural Network Acceleration
- Model Pruning Based on Quantified Similarity of Feature Maps
- NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
- A Survey of Coded Distributed Computing
- A Programmable Approach to Neural Network Compression
- LightGNN: Simple Graph Neural Network for Recommendation
- LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression
- Learning credit assignment
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Sparse deep neural networks for modeling aluminum electrolysis dynamics
- LCA: Loss Change Allocation for Neural Network Training
- Robust Lottery Tickets for Pre-trained Language Models
- Can Subnetwork Structure be the Key to Out-of-Distribution Generalization?
- Blockwise Self-Attention for Long Document Understanding
- FreezeNet: Full Performance by Reduced Storage Costs
- Automatic Mixed-Precision Quantization Search of BERT
- Federated Learning for Energy Constrained IoT devices: A systematic mapping study
- Path homologies of deep feedforward networks
- The Search for Sparse, Robust Neural Networks
- Mixed-Precision Embedding Using a Cache
- Evolving and Merging Hebbian Learning Rules: Increasing Generalization by Decreasing the Number of Rules
- VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation
- Achieving Adversarial Robustness via Sparsity
- Self-Damaging Contrastive Learning
- Bringing Giant Neural Networks Down to Earth with Unlabeled Data
- DASS: Differentiable Architecture Search for Sparse neural networks
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances
- A hybrid inference system for improved curvature estimation in the level-set method using machine learning
- ReFine: Re-randomization before Fine-tuning for Cross-domain Few-shot Learning
- Observation Space Matters: Benchmark and Optimization Algorithm
- Neural network relief: a pruning algorithm based on neural activity
- Fair Comparison: Quantifying Variance in Resultsfor Fine-grained Visual Categorization
- Layer Adaptive Node Selection in Bayesian Neural Networks: Statistical Guarantees and Implementation Details
- Towards Low-Latency Energy-Efficient Deep SNNs via Attention-Guided Compression
- Next2You: Robust Copresence Detection Based on Channel State Information
- Convolutional Neural Networks for Speech Controlled Prosthetic Hands
- Distilling Knowledge via Knowledge Review
- Sparsely Activated Networks
- Language Modeling on a SpiNNaker 2 Neuromorphic Chip
- Exploring Lottery Ticket Hypothesis in Media Recommender Systems
- Rethinking FUN: Frequency-Domain Utilization Networks
- Sparse Weight Activation Training
- HEMP: High-order Entropy Minimization for neural network comPression
- A Unified Paths Perspective for Pruning at Initialization
- Dynamic Multi-Branch Layers for On-Device Neural Machine Translation
- Neural Architecture Search using Property Guided Synthesis
- Campfire: Compressible, Regularization-Free, Structured Sparse Training for Hardware Accelerators
- Supermasks in Superposition
- SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference
- MARF: The Medial Atom Ray Field Object Representation
- Breaking the Activation Function Bottleneck through Adaptive Parameterization
- Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes
- MIME: Adapting a Single Neural Network for Multi-task Inference with Memory-efficient Dynamic Pruning
- Powerpropagation: A sparsity inducing weight reparameterisation
- Composition of Saliency Metrics for Channel Pruning with a Myopic Oracle
- Manipulating SGD with Data Ordering Attacks
- Calibrate and Prune: Improving Reliability of Lottery Tickets Through Prediction Calibration
- Patient trajectory prediction in the Mimic-III dataset, challenges and pitfalls
- Enabling Retrain-free Deep Neural Network Pruning using Surrogate Lagrangian Relaxation
- Alpha Discovery Neural Network based on Prior Knowledge
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- A Thorough Performance Benchmarking on Lightweight Embedding-based Recommender Systems
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- Collaboration of Experts: Achieving 80% Top-1 Accuracy on ImageNet with 100M FLOPs
- Adversarial Speaker Distillation for Countermeasure Model on Automatic Speaker Verification
- Adversarial Robustness through the Lens of Convolutional Filters
- Towards Practical Few-shot Federated NLP
- Model Compression with Two-stage Multi-teacher Knowledge Distillation for Web Question Answering System
- Conditioning of Random Feature Matrices: Double Descent and Generalization Error
- Shapley Value as Principled Metric for Structured Network Pruning
- Sparse evolutionary Deep Learning with over one million artificial neurons on commodity hardware
- Consistent Sparse Deep Learning: Theory and Computation
- GradSign: Model Performance Inference with Theoretical Insights
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- Intrinsically Sparse Long Short-Term Memory Networks
- KinePose: A temporally optimized inverse kinematics technique for 6DOF human pose estimation with biomechanical constraints
- Personalized Federated Learning by Structured and Unstructured Pruning under Data Heterogeneity
- Edge Bias in Federated Learning and its Solution by Buffered Knowledge Distillation
- Pufferfish: Communication-efficient Models At No Extra Cost
- Proof-of-Learning: Definitions and Practice
- Channel Equilibrium Networks for Learning Deep Representation
- Artificial neural networks condensation: A strategy to facilitate adaption of machine learning in medical settings by reducing computational burden
- High-contrast "gaudy" images improve the training of deep neural network models of visual cortex
- A Winning Hand: Compressing Deep Networks Can Improve Out-Of-Distribution Robustness
- Topological Insights into Sparse Neural Networks
- DeepLight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving
- COPS: Controlled Pruning Before Training Starts
- The Representation Theory of Neural Networks
- Hidden-Fold Networks: Random Recurrent Residuals Using Sparse Supermasks
- The Differentially Private Lottery Ticket Mechanism
- Training Neural Networks with Fixed Sparse Masks
- A Generalized Lottery Ticket Hypothesis
- Delta Keyword Transformer: Bringing Transformers to the Edge through Dynamically Pruned Multi-Head Self-Attention
- On Iterative Neural Network Pruning, Reinitialization, and the Similarity of Masks
- The Rediscovery Hypothesis: Language Models Need to Meet Linguistics
- Trainless Model Performance Estimation for Neural Architecture Search
- Data-Independent Structured Pruning of Neural Networks via Coresets
- SparseDNN: Fast Sparse Deep Learning Inference on CPUs
- Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
- Analysis of Gene Interaction Graphs as Prior Knowledge for Machine Learning Models
- Data-Model-Circuit Tri-Design for Ultra-Light Video Intelligence on Edge Devices
- ExplainFix: Explainable Spatially Fixed Deep Networks
- Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough
- Regularity Normalization: Neuroscience-Inspired Unsupervised Attention across Neural Network Layers
- Wider Networks Learn Better Features
- LNPT: Label-free Network Pruning and Training
- Active Subspace of Neural Networks: Structural Analysis and Universal Attacks
- DCI-ES: An Extended Disentanglement Framework with Connections to Identifiability
- A Practical Sparse Approximation for Real Time Recurrent Learning
- Pruning Filter in Filter
- Drawing Robust Scratch Tickets: Subnetworks with Inborn Robustness Are Found within Randomly Initialized Networks
- SASL: Saliency-Adaptive Sparsity Learning for Neural Network Acceleration
- Identifying Critical Neurons in ANN Architectures using Mixed Integer Programming
- Nonconvex sparse regularization for deep neural networks and its optimality
- Intelligence plays dice: Stochasticity is essential for machine learning
- Feature Products Yield Efficient Networks
- Learning Sparse & Ternary Neural Networks with Entropy-Constrained Trained Ternarization (EC2T)
- Properties Of Winning Tickets On Skin Lesion Classification
- Activation Density driven Energy-Efficient Pruning in Training
- Reconstructing cellular automata rules from observations at nonconsecutive times
- DECORE: Deep Compression with Reinforcement Learning
- Active multi-fidelity Bayesian online changepoint detection
- Learning to Act through Evolution of Neural Diversity in Random Neural Networks
- Train Flat, Then Compress: Sharpness-Aware Minimization Learns More Compressible Models
- Structured Compression by Weight Encryption for Unstructured Pruning and Quantization
- Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble Distillation
- A Conceptual Framework for Lifelong Learning
- Fixing the Teacher-Student Knowledge Discrepancy in Distillation
- Neural Network-based Automatic Factor Construction
- Testing the Genomic Bottleneck Hypothesis in Hebbian Meta-Learning
- Training Sparse Neural Networks using Compressed Sensing
- Deeplite Neutrino: An End-to-End Framework for Constrained Deep Learning Model Optimization
- Neural Networks at a Fraction with Pruned Quaternions
- SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions
- A Bregman Learning Framework for Sparse Neural Networks
- BWCP: Probabilistic Learning-to-Prune Channels for ConvNets via Batch Whitening
- EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models
- Deep network as memory space: complexity, generalization, disentangled representation and interpretability
- WaLDORf: Wasteless Language-model Distillation On Reading-comprehension
- Sparse Meta Networks for Sequential Adaptation and its Application to Adaptive Language Modelling
- Ultra-light deep MIR by trimming lottery tickets
- The curious case of developmental BERTology: On sparsity, transfer learning, generalization and the brain
- Convolutional neural networks compression with low rank and sparse tensor decompositions
- Self-Reorganizing and Rejuvenating CNNs for Increasing Model Capacity Utilization
- Sparsifying networks by traversing Geodesics
- Induction, Popper, and machine learning
- Training for temporal sparsity in deep neural networks, application in video processing
- Correlation Analysis between the Robustness of Sparse Neural Networks and their Random Hidden Structural Priors
- Generalizing Nucleus Recognition Model in Multi-source Images via Pruning
- Efficient Micro-Structured Weight Unification and Pruning for Neural Network Compression
- A Low-Compexity Deep Learning Framework For Acoustic Scene Classification
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
- MONCAE: Multi-Objective Neuroevolution of Convolutional Autoencoders
- Livewired Neural Networks: Making Neurons That Fire Together Wire Together
- Self-building Neural Networks
- MOFHEI: Model Optimizing Framework for Fast and Efficient Homomorphically Encrypted Neural Network Inference
- Dynamic Black-box Backdoor Attacks on IoT Sensory Data
- Inferring Thunderstorm Occurrence from Vertical Profiles of Convection-Permitting Simulations: Physical Insights from a Physical Deep Learning Model
- On Neural Networks as Infinite Tree-Structured Probabilistic Graphical Models
- On Lottery Tickets and Minimal Task Representations in Deep Reinforcement Learning
- Studying the Consistency and Composability of Lottery Ticket Pruning Masks
- Effect of the initial configuration of weights on the training and function of artificial neural networks
- The Heterogeneity Hypothesis: Finding Layer-Wise Differentiated Network Architectures
- SIPA: A Simple Framework for Efficient Networks
- Prune2Edge: A Multi-Phase Pruning Pipelines to Deep Ensemble Learning in IIoT
- A New MRAM-based Process In-Memory Accelerator for Efficient Neural Network Training with Floating Point Precision
- Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD
- Group Pruning using a Bounded-Lp norm for Group Gating and Regularization
- Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource Utilization
- Dense neural networks as sparse graphs and the lightning initialization
- Understanding the wiring evolution in differentiable neural architecture search
- Softer Pruning, Incremental Regularization
- Structured Ensembles: an Approach to Reduce the Memory Footprint of Ensemble Methods
- Learning Pruned Structure and Weights Simultaneously from Scratch: an Attention based Approach
- Zero-shot generalization using cascaded system-representations
- Is Feature Diversity Necessary in Neural Network Initialization?
- Sparsely Activated Networks: A new method for decomposing and compressing data
- Toward Compact Parameter Representations for Architecture-Agnostic Neural Network Compression
- Optimizing Connectivity through Network Gradients for Restricted Boltzmann Machines
- Does a sparse ReLU network training problem always admit an optimum?
- Geometric algorithms for predicting resilience and recovering damage in neural networks
- Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
- Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
- Polynomially Over-Parameterized Convolutional Neural Networks Contain Structured Strong Winning Lottery Tickets
- Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
- Optimization of DNN-based HSI Segmentation FPGA-based SoC for ADS: A Practical Approach
- Induced Feature Selection by Structured Pruning
- Convolutional Dictionary Learning in Hierarchical Networks
- Learning Digital Circuits: A Journey Through Weight Invariant Self-Pruning Neural Networks
- deepstruct -- linking deep learning and graph theory
- Data-dependent Pruning to find the Winning Lottery Ticket
- Neural Architecture Search via Bregman Iterations
- Solving hybrid machine learning tasks by traversing weight space geodesics
- Blending Pruning Criteria for Convolutional Neural Networks
- Deep learning for bioimage analysis
- Learning Compatible Embeddings
- Membership Inference Attacks on Lottery Ticket Networks
- Deconstructing the Structure of Sparse Neural Networks
- Interpreting Neural Networks as Gradual Argumentation Frameworks (Including Proof Appendix)
- A Quadratic Actor Network for Model-Free Reinforcement Learning
- Self-Constructing Neural Networks Through Random Mutation
- Network Adjustment: Channel Search Guided by FLOPs Utilization Ratio
- Data-driven Regularization via Racecar Training for Generalizing Neural Networks
- Neural networks adapting to datasets: learning network size and topology
- Out-of-the-box channel pruned networks
- CAZSL: Zero-Shot Regression for Pushing Models by Generalizing Through Context
- Neural Path Features and Neural Path Kernel : Understanding the role of gates in deep learning
- HALO: Learning to Prune Neural Networks with Shrinkage
- Against Membership Inference Attack: Pruning is All You Need