Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
arXiv:1712.06567
Abstract
Deep artificial neural networks (DNNs) are typically trained via gradient-based learning algorithms, namely backpropagation. Evolution strategies (ES) can rival backprop-based algorithms such as Q-learning and policy gradients on challenging deep reinforcement learning (RL) problems. However, ES can be considered a gradient-based algorithm because it performs stochastic gradient descent via an operation similar to a finite-difference approximation of the gradient. That raises the question of whether non-gradient-based evolutionary algorithms can work at DNN scales. Here we demonstrate they can: we evolve the weights of a DNN with a simple, gradient-free, population-based genetic algorithm (GA) and it performs well on hard deep RL problems, including Atari and humanoid locomotion. The Deep GA successfully evolves networks with over four million free parameters, the largest neural networks ever evolved with a traditional evolutionary algorithm. These results (1) expand our sense of the scale at which GAs can operate, (2) suggest intriguingly that in some cases following the gradient is not the best choice for optimizing performance, and (3) make immediately available the multitude of neuroevolution techniques that improve performance. We demonstrate the latter by showing that combining DNNs with novelty search, which encourages exploration on tasks with deceptive or sparse reward functions, can solve a high-dimensional problem on which reward-maximizing algorithms (e.g.\ DQN, A3C, ES, and the GA) fail. Additionally, the Deep GA is faster than ES, A3C, and DQN (it can train Atari in ${\raise.17ex\hbox{$\scriptstyle\sim$}}$4 hours on one desktop or ${\raise.17ex\hbox{$\scriptstyle\sim$}}$1 hour distributed on 720 cores), and enables a state-of-the-art, up to 10,000-fold compact encoding technique.
References in corpus (10)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Robots that can adapt like animals
- Self-Normalizing Neural Networks
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- Illuminating search spaces by mapping elites
- Massively Parallel Methods for Deep Reinforcement Learning
- Parameter Space Noise for Exploration
- A Distributional Perspective on Reinforcement Learning
- Hierarchical Representations for Efficient Architecture Search
- On the saddle point problem for non-convex optimization
Cited by in corpus (118)
- Neural Architecture Search: A Survey
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Neural network models and deep learning - a primer for biologists
- Exploration in Deep Reinforcement Learning: A Survey
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Go-Explore: a New Approach for Hard-Exploration Problems
- Neuroevolution in Deep Neural Networks: Current Trends and Future Challenges
- Ensemble Kalman Inversion: A Derivative-Free Technique For Machine Learning Tasks
- Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- Review: Deep Learning in Electron Microscopy
- Variational Quantum Reinforcement Learning via Evolutionary Optimization
- CEM-RL: Combining evolutionary and gradient-based methods for policy search
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- Proximal Distilled Evolutionary Reinforcement Learning
- GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
- Autocurricula and the Emergence of Innovation from Social Interaction: A Manifesto for Multi-Agent Intelligence Research
- MaskConnect: Connectivity Learning by Gradient Descent
- ES-MAML: Simple Hessian-Free Meta Learning
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
- Gradient Descent based Optimization Algorithms for Deep Learning Models Training
- Effective Diversity in Population Based Reinforcement Learning
- The Evolution of Reinforcement Learning in Quantitative Finance: A Survey
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference
- Autostacker: A Compositional Evolutionary Learning System
- Policy Manifold Search: Exploring the Manifold Hypothesis for Diversity-based Neuroevolution
- Evolving Modular Soft Robots without Explicit Inter-Module Communication using Local Self-Attention
- GADAM: Genetic-Evolutionary ADAM for Deep Neural Network Optimization
- GenAttack: Practical Black-box Attacks with Gradient-Free Optimization
- Evolutionary reinforcement learning of dynamical large deviations
- Learning Self-Imitating Diverse Policies
- Behavioural Repertoire via Generative Adversarial Policy Networks
- Correspondence between neuroevolution and gradient descent
- Empirical analysis of PGA-MAP-Elites for Neuroevolution in Uncertain Domains
- Black-box adversarial attacks using Evolution Strategies
- EvoJAX: Hardware-Accelerated Neuroevolution
- GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning
- Optimal quantum control via genetic algorithms for quantum state engineering in driven-resonator mediated networks
- Balancing Profit, Risk, and Sustainability for Portfolio Management
- Never look back - A modified EnKF method and its application to the training of neural networks without back propagation
- Playing Atari with Six Neurons
- Global convergence of neuron birth-death dynamics
- Evolutionary-Neural Hybrid Agents for Architecture Search
- Policy Evaluation Networks
- Inverse Design of a Graphene-Based Quantum Transducer via Neuroevolution
- Qualities, challenges and future of genetic algorithms: a literature review
- How You Act Tells a Lot: Privacy-Leakage Attack on Deep Reinforcement Learning
- Evolving the Behavior of Machines: From Micro to Macroevolution
- Solving the max-3-cut problem using synchronized dissipative networks
- Importance mixing: Improving sample reuse in evolutionary policy search methods
- Multi-objective Model-based Policy Search for Data-efficient Learning with Sparse Rewards
- Guiding Neuroevolution with Structural Objectives
- Neuroevolution of Neural Network Architectures Using CoDeepNEAT and Keras
- Deep Optimisation: Solving Combinatorial Optimisation Problems using Deep Neural Networks
- Diverse Auto-Curriculum is Critical for Successful Real-World Multiagent Learning Systems
- Evolvability ES: Scalable and Direct Optimization of Evolvability
- Smooth Exploration for Robotic Reinforcement Learning
- Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
- Evolving Inborn Knowledge For Fast Adaptation in Dynamic POMDP Problems
- Regenerating Soft Robots through Neural Cellular Automata
- Neuroevolution in Deep Learning: The Role of Neutrality
- Vulnerable road user detection: state-of-the-art and open challenges
- Policy Search with Rare Significant Events: Choosing the Right Partner to Cooperate with
- HyperNCA: Growing Developmental Networks with Neural Cellular Automata
- From Complexity to Simplicity: Adaptive ES-Active Subspaces for Blackbox Optimization
- Structured Control Nets for Deep Reinforcement Learning
- Symbolic Regression via Neural-Guided Genetic Programming Population Seeding
- Meta Learning Black-Box Population-Based Optimizers
- Interpretable discovery of new semiconductors with machine learning
- Hybrid Self-Attention NEAT: A novel evolutionary approach to improve the NEAT algorithm
- Fiber: A Platform for Efficient Development and Distributed Training for Reinforcement Learning and Population-Based Methods
- RSO: A Gradient Free Sampling Based Approach For Training Deep Neural Networks
- DLOPT: Deep Learning Optimization Library
- Simple Genetic Operators are Universal Approximators of Probability Distributions (and other Advantages of Expressive Encodings)
- A NEAT Quantum Error Decoder
- Evolutionary Augmentation Policy Optimization for Self-supervised Learning
- Learning Guidance Rewards with Trajectory-space Smoothing
- Sample-Efficient Automated Deep Reinforcement Learning
- The Adversarial Resilience Learning Architecture for AI-based Modelling, Exploration, and Operation of Complex Cyber-Physical Systems
- Neural Architecture Search Over a Graph Search Space
- A Continuous Optimisation Benchmark Suite from Neural Network Regression
- MDEA: Malware Detection with Evolutionary Adversarial Learning
- Evolutionary Construction of Convolutional Neural Networks
- Transferable Cost-Aware Security Policy Implementation for Malware Detection Using Deep Reinforcement Learning
- Evolving Self-supervised Neural Networks: Autonomous Intelligence from Evolved Self-teaching
- Multi-Path Policy Optimization
- Diverse Behavior Is What Game AI Needs: Generating Varied Human-Like Playing Styles Using Evolutionary Multi-Objective Deep Reinforcement Learning
- Learning Fitness Functions for Machine Programming
- Evolutionary Selective Imitation: Interpretable Agents by Imitation Learning Without a Demonstrator
- Optimizing Deep Neural Networks with Multiple Search Neuroevolution
- Competitiveness of MAP-Elites against Proximal Policy Optimization on locomotion tasks in deterministic simulations
- Switching Isotropic and Directional Exploration with Parameter Space Noise in Deep Reinforcement Learning
- A coevolutionary approach to deep multi-agent reinforcement learning
- Reducing catastrophic forgetting when evolving neural networks
- Monte-Carlo Tree Search for Policy Optimization
- Effects of Different Optimization Formulations in Evolutionary Reinforcement Learning on Diverse Behavior Generation
- Importance Weighted Evolution Strategies
- Accelerating Reinforcement Learning with a Directional-Gaussian-Smoothing Evolution Strategy
- Learning a Distributed Control Scheme for Demand Flexibility in Thermostatically Controlled Loads
- Generative Design of Hardware-aware DNNs
- Assessing and Accelerating Coverage in Deep Reinforcement Learning
- General Characterization of Agents by States they Visit
- Controllability, Multiplexing, and Transfer Learning in Networks using Evolutionary Learning
- SURREAL-System: Fully-Integrated Stack for Distributed Deep Reinforcement Learning
- ASCAI: Adaptive Sampling for acquiring Compact AI
- Recurrent Control Nets for Deep Reinforcement Learning
- VINE: An Open Source Interactive Data Visualization Tool for Neuroevolution
- Genetic-Gated Networks for Deep Reinforcement
- Why We Do Not Evolve Software? Analysis of Evolutionary Algorithms
- An Evolutionary Algorithm of Linear complexity: Application to Training of Deep Neural Networks
- Conditional Neural Architecture Search
- Reusability and Transferability of Macro Actions for Reinforcement Learning
- Supervising Unsupervised Learning with Evolutionary Algorithm in Deep Neural Network
- Evolving Indoor Navigational Strategies Using Gated Recurrent Units In NEAT
- ES-ENAS: Efficient Evolutionary Optimization for Large Hybrid Search Spaces
- The Foundations of Deep Learning with a Path Towards General Intelligence
- Population-based Global Optimisation Methods for Learning Long-term Dependencies with RNNs
- Quadratic speedup of global search using a biased crossover of two good solutions