Evolution Strategies as a Scalable Alternative to Reinforcement Learning
arXiv:1703.03864
Abstract
We explore the use of Evolution Strategies (ES), a class of black box optimization algorithms, as an alternative to popular MDP-based RL techniques such as Q-learning and Policy Gradients. Experiments on MuJoCo and Atari show that ES is a viable solution strategy that scales extremely well with the number of CPUs available: By using a novel communication strategy based on common random numbers, our ES implementation only needs to communicate scalars, making it possible to scale to over a thousand parallel workers. This allows us to solve 3D humanoid walking in 10 minutes and obtain competitive results on most Atari games after one hour of training. In addition, we highlight several advantages of ES as a black box optimization technique: it is invariant to action frequency and delayed rewards, tolerant of extremely long horizons, and does not need temporal discounting or value function approximation.
References in corpus (2)
Cited by in corpus (339)
- A Brief Survey of Deep Reinforcement Learning
- An Introduction to Deep Reinforcement Learning
- Noisy Networks for Exploration
- Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
- Exploration in Deep Reinforcement Learning: A Survey
- Adversarial Risk and the Dangers of Evaluating Against Weak Attacks
- Parameter Space Noise for Exploration
- Black-box Adversarial Attacks with Limited Queries and Information
- Ray: A Distributed Framework for Emerging AI Applications
- Agent57: Outperforming the Atari Human Benchmark
- Go-Explore: a New Approach for Hard-Exploration Problems
- Neuroevolution in Deep Neural Networks: Current Trends and Future Challenges
- Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
- Simple random search provides a competitive approach to reinforcement learning
- Evolved Policy Gradients
- Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
- Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents
- Deep Active Inference
- Neural Network-based Flight Control Systems: Present and Future
- CEM-RL: Combining evolutionary and gradient-based methods for policy search
- Backpropagation through the Void: Optimizing control variates for black-box gradient estimation
- Evolutionary Reinforcement Learning: A Survey
- Computer-inspired Quantum Experiments
- Global optimization of quantum dynamics with AlphaZero deep exploration
- Automated Evolutionary Approach for the Design of Composite Machine Learning Pipelines
- GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
- Reinforcement Learning for Integer Programming: Learning to Cut
- CHALET: Cornell House Agent Learning Environment
- DRLE: Decentralized Reinforcement Learning at the Edge for Traffic Light Control in the IoV
- MaskConnect: Connectivity Learning by Gradient Descent
- The Internet of Federated Things (IoFT): A Vision for the Future and In-depth Survey of Data-driven Approaches for Federated Learning
- Neuromorphic Hardware learns to learn
- Reducing Idleness in Financial Cloud Services via Multi-objective Evolutionary Reinforcement Learning based Load Balancer
- Learning the policy for mixed electric platoon control of automated and human-driven vehicles at signalized intersection: a random search approach
- ES-MAML: Simple Hessian-Free Meta Learning
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
- Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training Data
- NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural Networks
- Natural Evolutionary Strategies for Variational Quantum Computation
- Effective Diversity in Population Based Reinforcement Learning
- Enhanced POET: Open-Ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Diversity Policy Gradient for Sample Efficient Quality-Diversity Optimization
- Distilling Policy Distillation
- Gradientless Descent: High-Dimensional Zeroth-Order Optimization
- Curriculum-based Reinforcement Learning for Distribution System Critical Load Restoration
- Inverse Design of Grating Couplers Using the Policy Gradient Method from Reinforcement Learning
- Wasserstein Robust Reinforcement Learning
- MLGO: a Machine Learning Guided Compiler Optimizations Framework
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference
- Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree Search
- Zeroth-Order Regularized Optimization (ZORO): Approximately Sparse Gradients and Adaptive Sampling
- Optimal control of families of quantum gates
- AutoPhase: Juggling HLS Phase Orderings in Random Forests with Deep Reinforcement Learning
- Can Transfer Neuroevolution Tractably Solve Your Differential Equations?
- Policy Manifold Search: Exploring the Manifold Hypothesis for Diversity-based Neuroevolution
- Continuous-Discrete Reinforcement Learning for Hybrid Control in Robotics
- Evolving Modular Soft Robots without Explicit Inter-Module Communication using Local Self-Attention
- Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination
- Protocol Discovery for the Quantum Control of Majoranas by Differentiable Programming and Natural Evolution Strategies
- A directional Gaussian smoothing optimization method for computational inverse design in nanophotonics
- Query-Efficient Black-box Adversarial Examples (superceded)
- Robust Optimization through Neuroevolution
- Hessian-Aware Zeroth-Order Optimization for Black-Box Adversarial Attack
- Greedy Layerwise Learning Can Scale to ImageNet
- Evolutionary reinforcement learning of dynamical large deviations
- Autoencoding with a Classifier System
- Joint Multi-Dimension Pruning via Numerical Gradient Update
- Bayesian policy selection using active inference
- Learning Self-Imitating Diverse Policies
- Instance Weighted Incremental Evolution Strategies for Reinforcement Learning in Dynamic Environments
- Human-Level Reinforcement Learning through Theory-Based Modeling, Exploration, and Planning
- ZOOpt: Toolbox for Derivative-Free Optimization
- Correspondence between neuroevolution and gradient descent
- Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation
- A Tour of Reinforcement Learning: The View from Continuous Control
- AdvFlow: Inconspicuous Black-box Adversarial Attacks using Normalizing Flows
- Analysing Results from AI Benchmarks: Key Indicators and How to Obtain Them
- Reinforcement Learning Driven Heuristic Optimization
- A Primer on Zeroth-Order Optimization in Signal Processing and Machine Learning
- EvoJAX: Hardware-Accelerated Neuroevolution
- Limited Evaluation Cooperative Co-evolutionary Differential Evolution for Large-scale Neuroevolution
- From Pixels to Legs: Hierarchical Learning of Quadruped Locomotion
- Robotic Table Tennis: A Case Study into a High Speed Learning System
- Meta-SAC: Auto-tune the Entropy Temperature of Soft Actor-Critic via Metagradient
- Differential Variable Speed Limits Control for Freeway Recurrent Bottlenecks via Deep Reinforcement learning
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- Deep Curiosity Search: Intra-Life Exploration Can Improve Performance on Challenging Deep Reinforcement Learning Problems
- Sequence Modeling of Temporal Credit Assignment for Episodic Reinforcement Learning
- GPU-Accelerated Robotic Simulation for Distributed Reinforcement Learning
- Optimal quantum control via genetic algorithms for quantum state engineering in driven-resonator mediated networks
- On the Application of Danskin's Theorem to Derivative-Free Minimax Optimization
- Zeroth-Order Algorithms for Nonconvex Minimax Problems with Improved Complexities
- A primer on model-guided exploration of fitness landscapes for biological sequence design
- Global convergence of neuron birth-death dynamics
- Modern Deep Reinforcement Learning Algorithms
- Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
- On the Global Convergence of Imitation Learning: A Case for Linear Quadratic Regulator
- Stochastic Variational Optimization
- Playing Atari with Six Neurons
- Accelerated Deep Reinforcement Learning Based Load Shedding for Emergency Voltage Control
- Policy Evaluation Networks
- Evolving and Merging Hebbian Learning Rules: Increasing Generalization by Decreasing the Number of Rules
- Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement Learning
- Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves
- There are No Bit Parts for Sign Bits in Black-Box Attacks
- Efficient Convolutional Neural Network Training with Direct Feedback Alignment
- Zeroth-Order Algorithms for Smooth Saddle-Point Problems
- Mirror Descent Search and its Acceleration
- Value-aware Recommendation based on Reinforced Profit Maximization in E-commerce Systems
- Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model
- Evolving the Behavior of Machines: From Micro to Macroevolution
- Learning data augmentation policies using augmented random search
- Gradient-free Policy Architecture Search and Adaptation
- Learning To Simulate
- Convolutional Channel-wise Competitive Learning for the Forward-Forward Algorithm
- Lipizzaner: A System That Scales Robust Generative Adversarial Network Training
- Simplifying Deep Reinforcement Learning via Self-Supervision
- The Gap Between Model-Based and Model-Free Methods on the Linear Quadratic Regulator: An Asymptotic Viewpoint
- GeneCAI: Genetic Evolution for Acquiring Compact AI
- Multi-objective Model-based Policy Search for Data-efficient Learning with Sparse Rewards
- Rapidly Adaptable Legged Robots via Evolutionary Meta-Learning
- Importance mixing: Improving sample reuse in evolutionary policy search methods
- Neuroevolution of Neural Network Architectures Using CoDeepNEAT and Keras
- Provably Robust Blackbox Optimization for Reinforcement Learning
- The Role of Morphological Variation in Evolutionary Robotics: Maximizing Performance and Robustness
- Combining Model-Based and Model-Free Methods for Nonlinear Control: A Provably Convergent Policy Gradient Approach
- Prospects of reinforcement learning for the simultaneous damping of many mechanical modes
- Adversarial Imitation Learning via Random Search
- Guiding Neuroevolution with Structural Objectives
- Generalization Guarantees for Imitation Learning
- Dual Policy Distillation
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning
- Winning Isn't Everything: Enhancing Game Development with Intelligent Agents
- Map-based Multi-Policy Reinforcement Learning: Enhancing Adaptability of Robots by Deep Reinforcement Learning
- Safe Interactive Model-Based Learning
- Improving Gradient Estimation in Evolutionary Strategies With Past Descent Directions
- Maximum Mutation Reinforcement Learning for Scalable Control
- Limited-Memory Matrix Adaptation for Large Scale Black-box Optimization
- Evolvability ES: Scalable and Direct Optimization of Evolvability
- Black-box Adversarial Attacks on Video Recognition Models
- Learning Convex Optimization Control Policies
- Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
- Smooth Exploration for Robotic Reinforcement Learning
- An ODE Method to Prove the Geometric Convergence of Adaptive Stochastic Algorithms
- Comparing Deep Reinforcement Learning and Evolutionary Methods in Continuous Control
- Decision-based Universal Adversarial Attack
- Fast and Efficient Locomotion via Learned Gait Transitions
- Evolving Inborn Knowledge For Fast Adaptation in Dynamic POMDP Problems
- Automating Vehicles by Deep Reinforcement Learning using Task Separation with Hill Climbing
- On Entropy Regularized Path Integral Control for Trajectory Optimization
- Stochastic Subspace Descent
- Meta Learning Backpropagation And Improving It
- AdaDGS: An adaptive black-box optimization method with a nonlocal directional Gaussian smoothing gradient
- Patch-wise++ Perturbation for Adversarial Targeted Attacks
- Genetic Deep Learning for Lung Cancer Screening
- Linear interpolation gives better gradients than Gaussian smoothing in derivative-free optimization
- Policy Search with Rare Significant Events: Choosing the Right Partner to Cooperate with
- Data-Efficient Mutual Information Neural Estimator
- The Game of Tetris in Machine Learning
- ES-CTC: A Deep Neuroevolution Model for Cooperative Intelligent Freeway Traffic Control
- Efficiently avoiding saddle points with zero order methods: No gradients required
- Population-Guided Parallel Policy Search for Reinforcement Learning
- Improving Neural Network Training in Low Dimensional Random Bases
- Sample-Efficient Training of Robotic Guide Using Human Path Prediction Network
- HyperNCA: Growing Developmental Networks with Neural Cellular Automata
- Evolutionary optimisation of neural network models for fish collective behaviours in mixed groups of robots and zebrafish
- Dynamic control of self-assembly of quasicrystalline structures through reinforcement learning
- Online Hyper-parameter Tuning in Off-policy Learning via Evolutionary Strategies
- Adaptive Sampling Quasi-Newton Methods for Derivative-Free Stochastic Optimization
- Finding online neural update rules by learning to remember
- Learning Index Selection with Structured Action Spaces
- Structured Control Nets for Deep Reinforcement Learning
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- Adversarial Detection Avoidance Attacks: Evaluating the robustness of perceptual hashing-based client-side scanning
- Reinforcement Learning for Fair Dynamic Pricing
- Evolutionary Training and Abstraction Yields Algorithmic Generalization of Neural Computers
- From Complexity to Simplicity: Adaptive ES-Active Subspaces for Blackbox Optimization
- Calibration of Shared Equilibria in General Sum Partially Observable Markov Games
- Learning to Guide Random Search
- Transfer Learning versus Multi-agent Learning regarding Distributed Decision-Making in Highway Traffic
- Revenue, Relevance, Arbitrage and More: Joint Optimization Framework for Search Experiences in Two-Sided Marketplaces
- Learning Agile Locomotion via Adversarial Training
- Fiber: A Platform for Efficient Development and Distributed Training for Reinforcement Learning and Population-Based Methods
- When to be critical? Performance and evolvability in different regimes of neural Ising agents
- EvoGrad: Efficient Gradient-Based Meta-Learning and Hyperparameter Optimization
- Data-efficient Co-Adaptation of Morphology and Behaviour with Deep Reinforcement Learning
- Robotic Table Tennis with Model-Free Reinforcement Learning
- A Zeroth-Order Block Coordinate Descent Algorithm for Huge-Scale Black-Box Optimization
- Preventing Posterior Collapse with Levenshtein Variational Autoencoder
- Cooperative Coevolution for Non-Separable Large-Scale Black-Box Optimization: Convergence Analyses and Distributed Accelerations
- Extension of Direct Feedback Alignment to Convolutional and Recurrent Neural Network for Bio-plausible Deep Learning
- Hybrid Self-Attention NEAT: A novel evolutionary approach to improve the NEAT algorithm
- EGAN: Evolutional GAN for Ransomware Evasion
- Zero Shot Learning for Code Education: Rubric Sampling with Deep Learning Inference
- A Survey of Exploration Methods in Reinforcement Learning
- Combining Latent Space and Structured Kernels for Bayesian Optimization over Combinatorial Spaces
- Generalized Decision Transformer for Offline Hindsight Information Matching
- Meta Learning Black-Box Population-Based Optimizers
- A Stochastic Derivative Free Optimization Method with Momentum
- RSO: A Gradient Free Sampling Based Approach For Training Deep Neural Networks
- Taking Care of The Discretization Problem: A Comprehensive Study of the Discretization Problem and A Black-Box Adversarial Attack in Discrete Integer Domain
- DeepSearch: A Simple and Effective Blackbox Attack for Deep Neural Networks
- Learning to Score Behaviors for Guided Policy Optimization
- Trust-Region Variational Inference with Gaussian Mixture Models
- Adaptive Asynchronous Control Using Meta-learned Neural Ordinary Differential Equations
- A Comparison of Model-Free and Model Predictive Control for Price Responsive Water Heaters
- Distributed Evolution Strategies Using TPUs for Meta-Learning
- Learning Guidance Rewards with Trajectory-space Smoothing
- Improving Evolutionary Strategies with Generative Neural Networks
- AutoPhase: Compiler Phase-Ordering for High Level Synthesis with Deep Reinforcement Learning
- Sparse Stochastic Zeroth-Order Optimization with an Application to Bandit Structured Prediction
- On Machine Learning and Structure for Mobile Robots
- Learning Branching Heuristics for Propositional Model Counting
- Network of Evolvable Neural Units: Evolving to Learn at a Synaptic Level
- Epigenetic evolution of deep convolutional models
- Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation from a Blackbox Model
- Reward Tweaking: Maximizing the Total Reward While Planning for Short Horizons
- Policy Search by Target Distribution Learning for Continuous Control
- Recruitment-imitation Mechanism for Evolutionary Reinforcement Learning
- Augmented Random Search for Quadcopter Control: An alternative to Reinforcement Learning
- Harnessing Distribution Ratio Estimators for Learning Agents with Quality and Diversity
- Grid-Interactive Multi-Zone Building Control Using Reinforcement Learning with Global-Local Policy Search
- A Scalable Gradient-Free Method for Bayesian Experimental Design with Implicit Models
- Structured Monte Carlo Sampling for Nonisotropic Distributions via Determinantal Point Processes
- Evolutionary Variational Optimization of Generative Models
- Evolving Self-supervised Neural Networks: Autonomous Intelligence from Evolved Self-teaching
- Efficient Reinforcement Learning for StarCraft by Abstract Forward Models and Transfer Learning
- A Frank-Wolfe Framework for Efficient and Effective Adversarial Attacks
- Tuna: A Static Analysis Approach to Optimizing Deep Neural Networks
- Action Set Based Policy Optimization for Safe Power Grid Management
- Towards Automatic Actor-Critic Solutions to Continuous Control
- Finding mixed-strategy equilibria of continuous-action games without gradients using randomized policy networks
- Multi-Path Policy Optimization
- CoNES: Convex Natural Evolutionary Strategies
- Learning to Generate Synthetic 3D Training Data through Hybrid Gradient
- Information Bottleneck in Control Tasks with Recurrent Spiking Neural Networks
- Learning Fitness Functions for Machine Programming
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- MDEA: Malware Detection with Evolutionary Adversarial Learning
- Probably Approximately Correct Vision-Based Planning using Motion Primitives
- Variance Reduction for Evolution Strategies via Structured Control Variates
- Patch-wise Attack for Fooling Deep Neural Network
- Learning Synthetic Environments for Reinforcement Learning with Evolution Strategies
- Towards Query Efficient Black-box Attacks: An Input-free Perspective
- Hessian Inverse Approximation as Covariance for Random Perturbation in Black-Box Problems
- Testing the Genomic Bottleneck Hypothesis in Hebbian Meta-Learning
- Learning to Act through Evolution of Neural Diversity in Random Neural Networks
- Impression-Aware Recommender Systems
- Curiosity creates Diversity in Policy Search
- Leveraging Randomized Smoothing for Optimal Control of Nonsmooth Dynamical Systems
- Safe Exploration by Solving Early Terminated MDP
- Self-Evolutionary Optimization for Pareto Front Learning
- Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
- Unlocking Pixels for Reinforcement Learning via Implicit Attention
- Momentum Accelerates Evolutionary Dynamics
- Population-Based Evolution Optimizes a Meta-Learning Objective
- Zeroth-Order Supervised Policy Improvement
- Worm-level Control through Search-based Reinforcement Learning
- Novel Policy Seeking with Constrained Optimization
- A coevolutionary approach to deep multi-agent reinforcement learning
- MLE-guided parameter search for task loss minimization in neural sequence modeling
- Reinforcement Learning with Chromatic Networks for Compact Architecture Search
- Learning a Distributed Control Scheme for Demand Flexibility in Thermostatically Controlled Loads
- Evolutionary Selective Imitation: Interpretable Agents by Imitation Learning Without a Demonstrator
- Competitiveness of MAP-Elites against Proximal Policy Optimization on locomotion tasks in deterministic simulations
- A unified view of likelihood ratio and reparameterization gradients and an optimal importance sampling scheme
- Modelling zebrafish collective behaviours with multilayer perceptrons optimised by evolutionary algorithms
- Importance Weighted Evolution Strategies
- Monte-Carlo Tree Search for Policy Optimization
- Switching Isotropic and Directional Exploration with Parameter Space Noise in Deep Reinforcement Learning
- Neural Generative Models for Global Optimization with Gradients
- Accelerating Reinforcement Learning with a Directional-Gaussian-Smoothing Evolution Strategy
- Optimizing Deep Neural Networks with Multiple Search Neuroevolution
- Fractal AI: A fragile theory of intelligence
- Randomized Adversarial Imitation Learning for Autonomous Driving
- Genetic-Gated Networks for Deep Reinforcement
- Non-local Optimization: Imposing Structure on Optimization Problems by Relaxation
- State-Aware Variational Thompson Sampling for Deep Q-Networks
- Parameter-Based Value Functions
- Optimizing Sponsored Search Ranking Strategy by Deep Reinforcement Learning
- Adaptive Wind Driven Optimization Trained Artificial Neural Networks
- SURREAL-System: Fully-Integrated Stack for Distributed Deep Reinforcement Learning
- Gradients are Not All You Need
- Multi-Issue Bargaining With Deep Reinforcement Learning
- Minimal Neural Network Models for Permutation Invariant Agents
- On the Convergence of Prior-Guided Zeroth-Order Optimization Algorithms
- Sparse Perturbations for Improved Convergence in Stochastic Zeroth-Order Optimization
- Improved Learning in Evolution Strategies via Sparser Inter-Agent Network Topologies
- Evolution Strategies Converges to Finite Differences
- Autonomous Learning of Features for Control: Experiments with Embodied and Situated Agents
- Comparing heterogeneous entities using artificial neural networks of trainable weighted structural components and machine-learned activation functions
- Efficient Wasserstein Natural Gradients for Reinforcement Learning
- Analyzing Reinforcement Learning Benchmarks with Random Weight Guessing
- Recurrent Control Nets for Deep Reinforcement Learning
- Subgroup-based Rank-1 Lattice Quasi-Monte Carlo
- Learning Algorithmic Solutions to Symbolic Planning Tasks with a Neural Computer Architecture
- Periodic Intra-Ensemble Knowledge Distillation for Reinforcement Learning
- VINE: An Open Source Interactive Data Visualization Tool for Neuroevolution
- A General Framework on Enhancing Portfolio Management with Reinforcement Learning
- Questions to Guide the Future of Artificial Intelligence Research
- Learning in Sparse Rewards settings through Quality-Diversity algorithms
- Searching in the Forest for Local Bayesian Optimization
- Population-based Global Optimisation Methods for Learning Long-term Dependencies with RNNs
- Proximal Policy Optimization with Mixed Distributed Training
- Selecting for Selection: Learning To Balance Adaptive and Diversifying Pressures in Evolutionary Search
- FiDi-RL: Incorporating Deep Reinforcement Learning with Finite-Difference Policy Search for Efficient Learning of Continuous Control
- Reusability and Transferability of Macro Actions for Reinforcement Learning
- The Role of Modularity and Neuro-Regulation for the Production of Multiple Behaviors
- Meta Reinforcement Learning with Distribution of Exploration Parameters Learned by Evolution Strategies
- Black-box Optimizer with Implicit Natural Gradient
- Batch-Augmented Multi-Agent Reinforcement Learning for Efficient Traffic Signal Optimization
- Neural Auto-Curricula
- Distributionally Constrained Black-Box Stochastic Gradient Estimation and Optimization
- Mirror Natural Evolution Strategies
- Evolutionary Stochastic Policy Distillation
- Divergent Search for Few-Shot Image Classification
- Stochastic Flows and Geometric Optimization on the Orthogonal Group
- Learning to Win, Lose and Cooperate through Reward Signal Evolution
- Searching Search Spaces: Meta-evolving a Geometric Encoding for Neural Networks
- BOOK: Storing Algorithm-Invariant Episodes for Deep Reinforcement Learning
- Enabling Incremental Training with Forward Pass for Edge Devices
- How to Organize your Deep Reinforcement Learning Agents: The Importance of Communication Topology
- Improved Complexities for Stochastic Conditional Gradient Methods under Interpolation-like Conditions
- A Hybrid Gradient Method to Designing Bayesian Experiments for Implicit Models
- GitGraph - Architecture Search Space Creation through Frequent Computational Subgraph Mining
- Evolutionary Hyperparameter Optimization to Find Lightweight CNN Models for Autonomous Steering
- Can a Compact Neuronal Circuit Policy be Re-purposed to Learn Simple Robotic Control?
- GACEM: Generalized Autoregressive Cross Entropy Method for Multi-Modal Black Box Constraint Satisfaction
- Direct-Search for a Class of Stochastic Min-Max Problems
- Solving Atari Games Using Fractals And Entropy
- Neural Rate Control for Video Encoding using Imitation Learning
- Combine PPO with NES to Improve Exploration
- Escaping Saddle Points for Zeroth-order Nonconvex Optimization using Estimated Gradient Descent
- Training Efficiency and Robustness in Deep Learning
- Diverse Exploration via Conjugate Policies for Policy Gradient Methods
- Learning Deep Energy Shaping Policies for Stability-Guaranteed Manipulation
- On the Second-order Convergence Properties of Random Search Methods
- Ranking Cost: Building An Efficient and Scalable Circuit Routing Planner with Evolution-Based Optimization