Neural Architecture Search with Reinforcement Learning
arXiv:1611.01578
Abstract
Neural networks are powerful and flexible models that work well for many difficult learning tasks in image, speech and natural language understanding. Despite their success, neural networks are still hard to design. In this paper, we use a recurrent network to generate the model descriptions of neural networks and train this RNN with reinforcement learning to maximize the expected accuracy of the generated architectures on a validation set. On the CIFAR-10 dataset, our method, starting from scratch, can design a novel network architecture that rivals the best human-invented architecture in terms of test set accuracy. Our CIFAR-10 model achieves a test error rate of 3.65, which is 0.09 percent better and 1.05x faster than the previous state-of-the-art model that used a similar architectural scheme. On the Penn Treebank dataset, our model can compose a novel recurrent cell that outperforms the widely-used LSTM cell, and other state-of-the-art baselines. Our cell achieves a test set perplexity of 62.4 on the Penn Treebank, which is 3.6 perplexity better than the previous state-of-the-art model. The cell can also be transferred to the character language modeling task on PTB and achieves a state-of-the-art perplexity of 1.214.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Striving for Simplicity: The All Convolutional Net
- Recurrent Neural Network Regularization
- Theoretical Models of Learning to Learn
- Pointer Sentinel Mixture Models
- Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
Cited by in corpus (138)
- A Brief Survey of Deep Reinforcement Learning
- Searching for Activation Functions
- Regularizing and Optimizing LSTM Language Models
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- Robust Adversarial Reinforcement Learning
- Learning to reinforcement learn
- Neural Optimizer Search with Reinforcement Learning
- Simple And Efficient Architecture Search for Convolutional Neural Networks
- Automated Machine Learning: State-of-The-Art and Open Challenges
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- Optimization for deep learning: theory and algorithms
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning
- FasterSeg: Searching for Faster Real-time Semantic Segmentation
- AdaNet: Adaptive Structural Learning of Artificial Neural Networks
- Automated Machine Learning in Practice: State of the Art and Recent Results
- Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing
- Adversarial AutoAugment
- Peephole: Predicting Network Performance Before Training
- SpArSe: Sparse Architecture Search for CNNs on Resource-Constrained Microcontrollers
- Exploring Randomly Wired Neural Networks for Image Recognition
- Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic Forgetting
- Probabilistic Neural Architecture Search
- Learning to Continually Learn
- Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
- Dynamic Evaluation of Neural Sequence Models
- On the State of the Art of Evaluation in Neural Language Models
- Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training Data
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Revisiting Activation Regularization for Language RNNs
- Fast-Slow Recurrent Neural Networks
- AGAN: Towards Automated Design of Generative Adversarial Networks
- Multi-Objective Reinforced Evolution in Mobile Neural Architecture Search
- Hyp-RL : Hyperparameter Optimization by Reinforcement Learning
- Deeper Insights into Weight Sharing in Neural Architecture Search
- MaxUp: A Simple Way to Improve Generalization of Neural Network Training
- Neural Graph Evolution: Towards Efficient Automatic Robot Design
- Computation Reallocation for Object Detection
- Recurrent Additive Networks
- Learning to Prune Filters in Convolutional Neural Networks
- Learning Graph Convolutional Network for Skeleton-based Human Action Recognition by Neural Searching
- RC-DARTS: Resource Constrained Differentiable Architecture Search
- Connectivity Learning in Multi-Branch Networks
- Deep Long Audio Inpainting
- A Comprehensive Survey of Multilingual Neural Machine Translation
- Intriguing Properties of Adversarial Examples
- XNAS: Neural Architecture Search with Expert Advice
- AutoEmb: Automated Embedding Dimensionality Search in Streaming Recommendations
- Not All Ops Are Created Equal!
- Reinforcement Learning Driven Heuristic Optimization
- HARK Side of Deep Learning -- From Grad Student Descent to Automated Machine Learning
- Question Answering from Unstructured Text by Retrieval and Comprehension
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation
- Exploring Benefits of Transfer Learning in Neural Machine Translation
- Real-time Federated Evolutionary Neural Architecture Search
- Large Scale Evolution of Convolutional Neural Networks Using Volunteer Computing
- Multigrid Predictive Filter Flow for Unsupervised Learning on Videos
- Using Small Proxy Datasets to Accelerate Hyperparameter Search
- Efficient Differentiable Neural Architecture Search with Meta Kernels
- Learning to Optimize in Swarms
- Convolution Neural Network Architecture Learning for Remote Sensing Scene Classification
- Learnable Embedding Space for Efficient Neural Architecture Compression
- FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge
- Reinforcement Learning and Adaptive Sampling for Optimized DNN Compilation
- Fine-Grained Neural Architecture Search
- BETANAS: BalancEd TrAining and selective drop for Neural Architecture Search
- Gradient-free Policy Architecture Search and Adaptation
- Dynamic Optimization of Neural Network Structures Using Probabilistic Modeling
- A Review of Meta-Reinforcement Learning for Deep Neural Networks Architecture Search
- Neuroevolution of Neural Network Architectures Using CoDeepNEAT and Keras
- DVOLVER: Efficient Pareto-Optimal Neural Network Architecture Search
- Bayesian Learning of Neural Network Architectures
- Meta-Learning for Contextual Bandit Exploration
- Map-based Multi-Policy Reinforcement Learning: Enhancing Adaptability of Robots by Deep Reinforcement Learning
- Nonparametric Neural Networks
- Techniques for Automated Machine Learning
- Automatic Model Selection for Neural Networks
- Improving Neural Architecture Search Image Classifiers via Ensemble Learning
- ProBO: Versatile Bayesian Optimization Using Any Probabilistic Programming Language
- Neural Architecture Search For Fault Diagnosis
- Learning to update Auto-associative Memory in Recurrent Neural Networks for Improving Sequence Memorization
- Data-driven Neural Architecture Learning For Financial Time-series Forecasting
- Accuracy vs. Efficiency: Achieving Both through FPGA-Implementation Aware Neural Architecture Search
- Automatically Searching for U-Net Image Translator Architecture
- Domain-Aware Dynamic Networks
- Multi-objective Neural Architecture Search via Non-stationary Policy Gradient
- Catalyst.RL: A Distributed Framework for Reproducible RL Research
- Towards Oracle Knowledge Distillation with Neural Architecture Search
- DARC: Differentiable ARchitecture Compression
- Transfer Learning to Learn with Multitask Neural Model Search
- Deep Neural Architecture Search with Deep Graph Bayesian Optimization
- Video Action Recognition Via Neural Architecture Searching
- TextNAS: A Neural Architecture Search Space tailored for Text Representation
- Evolving Deep Neural Networks by Multi-objective Particle Swarm Optimization for Image Classification
- Generalizable Resource Allocation in Stream Processing via Deep Reinforcement Learning
- RAPDARTS: Resource-Aware Progressive Differentiable Architecture Search
- Learning Less-Overlapping Representations
- EPNAS: Efficient Progressive Neural Architecture Search
- StyleNAS: An Empirical Study of Neural Architecture Search to Uncover Surprisingly Fast End-to-End Universal Style Transfer Networks
- A Generative Model for Sampling High-Performance and Diverse Weights for Neural Networks
- WeNet: Weighted Networks for Recurrent Network Architecture Search
- Regularize, Expand and Compress: Multi-task based Lifelong Learning via NonExpansive AutoML
- Capacity allocation analysis of neural networks: A tool for principled architecture design
- Learned Indexes for Dynamic Workloads
- RandomNet: Towards Fully Automatic Neural Architecture Design for Multimodal Learning
- Intelligence, physics and information -- the tradeoff between accuracy and simplicity in machine learning
- ResNetX: a more disordered and deeper network architecture
- Data-Driven Compression of Convolutional Neural Networks
- Exploiting Operation Importance for Differentiable Neural Architecture Search
- Explaining Transition Systems through Program Induction
- Automated Problem Identification: Regression vs Classification via Evolutionary Deep Networks
- DeepSwarm: Optimising Convolutional Neural Networks using Swarm Intelligence
- In folly ripe. In reason rotten. Putting machine theology to rest
- Using Program Induction to Interpret Transition System Dynamics
- URNet : User-Resizable Residual Networks with Conditional Gating Module
- ImmuNetNAS: An Immune-network approach for searching Convolutional Neural Network Architectures
- Learning to Cope with Adversarial Attacks
- Genetic Network Architecture Search
- Rotational Unit of Memory
- A Feasible Framework for Arbitrary-Shaped Scene Text Recognition
- MLFriend: Interactive Prediction Task Recommendation for Event-Driven Time-Series Data
- Design of Artificial Intelligence Agents for Games using Deep Reinforcement Learning
- Mise en abyme with artificial intelligence: how to predict the accuracy of NN, applied to hyper-parameter tuning
- Automated Deep Abstractions for Stochastic Chemical Reaction Networks
- Neural Inheritance Relation Guided One-Shot Layer Assignment Search
- ADWPNAS: Architecture-Driven Weight Prediction for Neural Architecture Search
- On the potential for open-endedness in neural networks
- -LBI: Stochastic Split Linearized Bregman Iterations for Parsimonious Deep Learning
- Inferring Mesoscale Models of Neural Computation
- Photofeeler-D3: A Neural Network with Voter Modeling for Dating Photo Impression Prediction
- Seeing Convolution Through the Eyes of Finite Transformation Semigroup Theory: An Abstract Algebraic Interpretation of Convolutional Neural Networks
- AmoebaContact and GDFold: a new pipeline for rapid prediction of protein structures
- ARMIN: Towards a More Efficient and Light-weight Recurrent Memory Network
- An Embedded Deep Learning based Word Prediction
- Entropy Non-increasing Games for the Improvement of Dataflow Programming
- Patterns versus Characters in Subword-aware Neural Language Modeling
- Input-to-Output Gate to Improve RNN Language Models
- Rethinking Layer-wise Feature Amounts in Convolutional Neural Network Architectures