Simple random search provides a competitive approach to reinforcement learning
arXiv:1803.07055
Abstract
A common belief in model-free reinforcement learning is that methods based on random search in the parameter space of policies exhibit significantly worse sample complexity than those that explore the space of actions. We dispel such beliefs by introducing a random search method for training static, linear policies for continuous control problems, matching state-of-the-art sample efficiency on the benchmark MuJoCo locomotion tasks. Our method also finds a nearly optimal controller for a challenging instance of the Linear Quadratic Regulator, a classical problem in control theory, when the dynamics are not known. Computationally, our random search algorithm is at least 15 times more efficient than the fastest competing model-free methods on these benchmarks. We take advantage of this computational efficiency to evaluate the performance of our method over hundreds of random seeds and many different hyperparameter configurations for each benchmark task. Our simulations highlight a high variability in performance in these benchmark tasks, suggesting that commonly used estimations of sample efficiency do not adequately evaluate the performance of RL algorithms.
22 pages, 5 figures, 9 tables
References in corpus (16)
- Adam: A Method for Stochastic Optimization
- Continuous control with deep reinforcement learning
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Emergence of Locomotion Behaviours in Rich Environments
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- Parameter Space Noise for Exploration
- Ray: A Distributed Framework for Emerging AI Applications
- Sample Efficient Actor-Critic with Experience Replay
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
- On the Sample Complexity of the Linear Quadratic Regulator
- Least-Squares Temporal Difference Learning for the Linear Quadratic Regulator
- Towards Generalization and Simplicity in Continuous Control
Cited by in corpus (86)
- AutoAugment: Learning Augmentation Policies from Data
- Reward Constrained Policy Optimization
- Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- Neural Network-based Flight Control Systems: Present and Future
- Global optimization of quantum dynamics with AlphaZero deep exploration
- DRLE: Decentralized Reinforcement Learning at the Edge for Traffic Light Control in the IoV
- How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics
- ES-MAML: Simple Hessian-Free Meta Learning
- Effective Diversity in Population Based Reinforcement Learning
- Learning to Locomote: Understanding How Environment Design Matters for Deep Reinforcement Learning
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Gradientless Descent: High-Dimensional Zeroth-Order Optimization
- Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search
- Zeroth-Order Regularized Optimization (ZORO): Approximately Sparse Gradients and Adaptive Sampling
- The Scientific Method in the Science of Machine Learning
- A Closer Look at Deep Policy Gradients
- Finite-time Analysis of Approximate Policy Iteration for the Linear Quadratic Regulator
- How to pick the domain randomization parameters for sim-to-real transfer of reinforcement learning policies?
- A Tour of Reinforcement Learning: The View from Continuous Control
- Boosting Trust Region Policy Optimization by Normalizing Flows Policy
- S-RL Toolbox: Environments, Datasets and Evaluation Metrics for State Representation Learning
- Observational Overfitting in Reinforcement Learning
- GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning
- Accelerated Deep Reinforcement Learning Based Load Shedding for Emergency Voltage Control
- Policy Evaluation Networks
- Discretizing Continuous Action Space for On-Policy Optimization
- SADA: Semantic Adversarial Diagnostic Attacks for Autonomous Applications
- Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video
- Deep Reinforcement Learning Designed Shinnar-Le Roux RF Pulse using Root-Flipping: DeepRF_SLR
- Learning data augmentation policies using augmented random search
- The Gap Between Model-Based and Model-Free Methods on the Linear Quadratic Regulator: An Asymptotic Viewpoint
- Learning Fast Adaptation with Meta Strategy Optimization
- Provably Robust Blackbox Optimization for Reinforcement Learning
- Adversarial Imitation Learning via Random Search
- Feature Partitioning for Efficient Multi-Task Architectures
- Rapidly Adaptable Legged Robots via Evolutionary Meta-Learning
- A One-bit, Comparison-Based Gradient Estimator
- Constrained Reinforcement Learning Has Zero Duality Gap
- Hierarchical Reinforcement Learning for Quadruped Locomotion
- Smooth Exploration for Robotic Reinforcement Learning
- Structured Neural Network Dynamics for Model-based Control
- Online Hyper-parameter Tuning in Off-policy Learning via Evolutionary Strategies
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- From Complexity to Simplicity: Adaptive ES-Active Subspaces for Blackbox Optimization
- Learning Index Selection with Structured Action Spaces
- Implicit Policy for Reinforcement Learning
- Robotic Table Tennis with Model-Free Reinforcement Learning
- Learning Agile Locomotion via Adversarial Training
- How Much Do Unstated Problem Constraints Limit Deep Robotic Reinforcement Learning?
- Dynamics and Domain Randomized Gait Modulation with Bezier Curves for Sim-to-Real Legged Locomotion
- A Stochastic Derivative Free Optimization Method with Momentum
- Sparse Stochastic Zeroth-Order Optimization with an Application to Bandit Structured Prediction
- Learning Long-Term Reward Redistribution via Randomized Return Decomposition
- Reward Tweaking: Maximizing the Total Reward While Planning for Short Horizons
- CoNES: Convex Natural Evolutionary Strategies
- Post-Estimation Smoothing: A Simple Baseline for Learning with Side Information
- Augmented Random Search for Quadcopter Control: An alternative to Reinforcement Learning
- Adversarial Imitation Learning via Random Search in Lane Change Decision-Making
- Hessian Inverse Approximation as Covariance for Random Perturbation in Black-Box Problems
- Structured Monte Carlo Sampling for Nonisotropic Distributions via Determinantal Point Processes
- Policy Search by Target Distribution Learning for Continuous Control
- Learning a Distributed Control Scheme for Demand Flexibility in Thermostatically Controlled Loads
- MLE-guided parameter search for task loss minimization in neural sequence modeling
- Zeroth-Order Supervised Policy Improvement
- Reinforcement Learning with Chromatic Networks for Compact Architecture Search
- RLgraph: Modular Computation Graphs for Deep Reinforcement Learning
- Randomized Adversarial Imitation Learning for Autonomous Driving
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy Evaluation
- Extracting Latent State Representations with Linear Dynamics from Rich Observations
- Sparse Perturbations for Improved Convergence in Stochastic Zeroth-Order Optimization
- Learning Agile Locomotion Skills with a Mentor
- Can a Compact Neuronal Circuit Policy be Re-purposed to Learn Simple Robotic Control?
- On the Second-order Convergence Properties of Random Search Methods
- Exploration in Action Space
- Randomized Policy Learning for Continuous State and Action MDPs
- Wield: Systematic Reinforcement Learning With Progressive Randomization
- ES-ENAS: Efficient Evolutionary Optimization for Large Hybrid Search Spaces
- Explaining Inference Queries with Bayesian Optimization
- Parameter Critic: a Model Free Variance Reduction Method Through Imperishable Samples
- Depth and nonlinearity induce implicit exploration for RL
- Accelerating Optimization and Reinforcement Learning with Quasi-Stochastic Approximation
- All-Action Policy Gradient Methods: A Numerical Integration Approach
- Policy Search using Dynamic Mirror Descent MPC for Model Free Off Policy RL