Net2Net: Accelerating Learning via Knowledge Transfer
arXiv:1511.05641
Abstract
We introduce techniques for rapidly transferring the information stored in one neural net into another neural net. The main purpose is to accelerate the training of a significantly larger neural net. During real-world workflows, one often trains very many different neural networks during the experimentation and design process. This is a wasteful process in which each new model is trained from scratch. Our Net2Net technique accelerates the experimentation process by instantaneously transferring the knowledge from a previous network to each new deeper or wider network. Our techniques are based on the concept of function-preserving transformations between neural network specifications. This differs from previous approaches to pre-training that altered the function represented by a neural net when adding layers to it. Using our knowledge transfer mechanism to add depth to Inception modules, we demonstrate a new state of the art accuracy rating on the ImageNet dataset.
ICLR 2016 submission
References in corpus (3)
Cited by in corpus (69)
- Knowledge Distillation: A Survey
- Geometric deep learning: going beyond Euclidean data
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- A Survey on Evolutionary Neural Architecture Search
- PathNet: Evolution Channels Gradient Descent in Super Neural Networks
- DynGEM: Deep Embedding Method for Dynamic Graphs
- A Comprehensive Survey of Neural Architecture Search: Challenges and Solutions
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
- Universal representations:The missing link between faces, text, planktons, and cat breeds
- LAMOL: LAnguage MOdeling for Lifelong Language Learning
- Practical Blind Membership Inference Attack via Differential Comparisons
- Best of Both Worlds: Transferring Knowledge from Discriminative Learning to a Generative Visual Dialog Model
- Knowledge Distillation in Generations: More Tolerant Teachers Educate Better Students
- Continual Learning in Sensor-based Human Activity Recognition: an Empirical Benchmark Analysis
- Delving Deeper into MOOC Student Dropout Prediction
- Language Models with Transformers
- Alzheimer's Disease Diagnostics by a Deeply Supervised Adaptable 3D Convolutional Network
- Noisy Differentiable Architecture Search
- LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5
- Forward Thinking: Building and Training Neural Networks One Layer at a Time
- NeurObfuscator: A Full-stack Obfuscation Tool to Mitigate Neural Architecture Stealing
- Lifelong Adaptive Machine Learning for Sensor-based Human Activity Recognition Using Prototypical Networks
- StackRec: Efficient Training of Very Deep Sequential Recommender Models by Iterative Stacking
- Progressive Reinforcement Learning with Distillation for Multi-Skilled Motion Control
- Distilling Knowledge from Graph Convolutional Networks
- Multiple Expert Brainstorming for Domain Adaptive Person Re-identification
- Bag of Baselines for Multi-objective Joint Neural Architecture Search and Hyperparameter Optimization
- Towards Training Recurrent Neural Networks for Lifelong Learning
- Constructing Deep Neural Networks by Bayesian Network Structure Learning
- Training multi-objective/multi-task collocation physics-informed neural network with student/teachers transfer learnings
- Anomaly Detection through Transfer Learning in Agriculture and Manufacturing IoT Systems
- OpenEI: An Open Framework for Edge Intelligence
- Modularized Morphing of Neural Networks
- A Framework for Searching for General Artificial Intelligence
- Multiresolution Convolutional Autoencoders
- Representation Stability as a Regularizer for Improved Text Analytics Transfer Learning
- Towards Self-Adaptive Metric Learning On the Fly
- Learning Deep Representations with Probabilistic Knowledge Transfer
- Neural Architecture Search using Deep Neural Networks and Monte Carlo Tree Search
- Incremental Learning Using a Grow-and-Prune Paradigm with Efficient Neural Networks
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- Beyond Fine Tuning: A Modular Approach to Learning on Small Data
- Reducing the Training Time of Neural Networks by Partitioning
- Parallelizing Over Artificial Neural Network Training Runs with Multigrid
- AutoGrow: Automatic Layer Growing in Deep Convolutional Networks
- BI-MAML: Balanced Incremental Approach for Meta Learning
- Unsupervised Deep Feature Transfer for Low Resolution Image Classification
- Numerical Matrix Decomposition
- Intra-Ensemble in Neural Networks
- Simplified Stochastic Feedforward Neural Networks
- Regularize, Expand and Compress: Multi-task based Lifelong Learning via NonExpansive AutoML
- A novel method for identifying the deep neural network model with the Serial Number
- Hyperparameter Transfer Across Developer Adjustments
- Distilling Pixel-Wise Feature Similarities for Semantic Segmentation
- CompNet: Neural networks growing via the compact network morphism
- The Elastic Lottery Ticket Hypothesis
- Efficient Model Performance Estimation via Feature Histories
- Greedy Network Enlarging
- AIPerf: Automated machine learning as an AI-HPC benchmark
- Neural Network Surgery with Sets
- Auto Deep Compression by Reinforcement Learning Based Actor-Critic Structure
- Temporal Graph Offset Reconstruction: Towards Temporally Robust Graph Representation Learning
- Learning Realistic Patterns from Unrealistic Stimuli: Generalization and Data Anonymization
- Transfer Learning Between Different Architectures Via Weights Injection
- An Improving Framework of regularization for Network Compression
- A Useful Motif for Flexible Task Learning in an Embodied Two-Dimensional Visual Environment
- General AI Challenge - Round One: Gradual Learning
- Modular Continual Learning in a Unified Visual Environment