Noisy Networks for Exploration
arXiv:1706.10295
Abstract
We introduce NoisyNet, a deep reinforcement learning agent with parametric noise added to its weights, and show that the induced stochasticity of the agent's policy can be used to aid efficient exploration. The parameters of the noise are learned with gradient descent along with the remaining network weights. NoisyNet is straightforward to implement and adds little computational overhead. We find that replacing the conventional exploration heuristics for A3C, DQN and dueling agents (entropy reward and -greedy respectively) with NoisyNet yields substantially higher scores for a wide range of Atari games, in some cases advancing the agent from sub to super-human performance.
ICLR 2018
References in corpus (8)
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Parameter Space Noise for Exploration
- Evolutionary Algorithms for Reinforcement Learning
- Count-Based Exploration with Neural Density Models
- Deep Exploration via Randomized Value Functions
- Bayesian Recurrent Neural Networks
- Kalman Temporal Differences
- Minimax Regret Bounds for Reinforcement Learning
Cited by in corpus (152)
- Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
- Exploration in Deep Reinforcement Learning: A Survey
- A Review of Deep Reinforcement Learning for Smart Building Energy Management
- Parameter Space Noise for Exploration
- Large-Scale Study of Curiosity-Driven Learning
- Model-Ensemble Trust-Region Policy Optimization
- Exploration by Random Network Distillation
- Implicit Quantile Networks for Distributional Reinforcement Learning
- Agent57: Outperforming the Atari Human Benchmark
- An Application of Deep Reinforcement Learning to Algorithmic Trading
- CURL: Contrastive Unsupervised Representations for Reinforcement Learning
- Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Meta-Reinforcement Learning of Structured Exploration Strategies
- Autonomous Unmanned Aerial Vehicle Navigation using Reinforcement Learning: A Systematic Review
- AI Safety Gridworlds
- Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents
- IG-RL: Inductive Graph Reinforcement Learning for Massive-Scale Traffic Signal Control
- Review: Deep Learning in Electron Microscopy
- RUDDER: Return Decomposition for Delayed Rewards
- Accelerated Methods for Deep Reinforcement Learning
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- A survey on intrinsic motivation in reinforcement learning
- Efficient Exploration via State Marginal Matching
- Towards V2I Age-aware Fairness Access: A DQN Based Intelligent Vehicular Node Training and Test Method
- Horizon: Facebook's Open Source Applied Reinforcement Learning Platform
- GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
- ChainerRL: A Deep Reinforcement Learning Library
- Deep Reinforcement Learning Control of Quantum Cartpoles
- A Deep Reinforcement Learning Approach for the Meal Delivery Problem
- Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control
- Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
- Deep reinforcement learning for guidewire navigation in coronary artery phantom
- Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?
- Data-Efficient Reinforcement Learning with Self-Predictive Representations
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
- Improving Generalization in Reinforcement Learning with Mixture Regularization
- Effective Diversity in Population Based Reinforcement Learning
- Bayesian Sequential Optimal Experimental Design for Nonlinear Models Using Policy Gradient Reinforcement Learning
- Whatever Does Not Kill Deep Reinforcement Learning, Makes It Stronger
- Noisy Differentiable Architecture Search
- Prioritized Sequence Experience Replay
- Neural Thompson Sampling
- The Evolution of Reinforcement Learning in Quantitative Finance: A Survey
- Designing Deep Reinforcement Learning for Human Parameter Exploration
- Mastering Atari with Discrete World Models
- Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
- Benchmarking Bonus-Based Exploration Methods on the Arcade Learning Environment
- Improving robot navigation in crowded environments using intrinsic rewards
- State Entropy Maximization with Random Encoders for Efficient Exploration
- Sample-Efficient Imitation Learning via Generative Adversarial Nets
- The Faults in Our Pi Stars: Security Issues and Open Challenges in Deep Reinforcement Learning
- Revisiting Rainbow: Promoting more Insightful and Inclusive Deep Reinforcement Learning Research
- The problem with DDPG: understanding failures in deterministic environments with sparse rewards
- Learning First-to-Spike Policies for Neuromorphic Control Using Policy Gradients
- Is Deep Reinforcement Learning Really Superhuman on Atari? Leveling the playing field
- BeBold: Exploration Beyond the Boundary of Explored Regions
- Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning
- Self-Supervised Exploration via Disagreement
- Applications of Deep Reinforcement Learning in Communications and Networking: A Survey
- Online Self-Supervised Learning for Object Picking: Detecting Optimum Grasping Position using a Metric Learning Approach
- Meta-learning curiosity algorithms
- Accelerating Reinforcement Learning through GPU Atari Emulation
- Randomized Value Functions via Multiplicative Normalizing Flows
- Worst-Case Regret Bounds for Exploration via Randomized Value Functions
- Variance Networks: When Expectation Does Not Meet Your Expectations
- Learn to Interpret Atari Agents
- CCLF: A Contrastive-Curiosity-Driven Learning Framework for Sample-Efficient Reinforcement Learning
- A Universal Adversarial Policy for Text Classifiers
- Spectral Normalisation for Deep Reinforcement Learning: an Optimisation Perspective
- Surprising Negative Results for Generative Adversarial Tree Search
- Reinforced Epidemic Control: Saving Both Lives and Economy
- Student-Initiated Action Advising via Advice Novelty
- Interactive Language Learning by Question Answering
- Exploration by Distributional Reinforcement Learning
- Reinforcement Learning-based Switching Controller for a Milliscale Robot in a Constrained Environment
- Smooth Exploration for Robotic Reinforcement Learning
- Long-Term Visitation Value for Deep Exploration in Sparse Reward Reinforcement Learning
- Variational Deep Q Network
- A review of motion planning algorithms for intelligent robotics
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning
- On the Complexity of Exploration in Goal-Driven Navigation
- Variational Adaptive-Newton Method for Explorative Learning
- Temporal Difference Uncertainties as a Signal for Exploration
- VFunc: a Deep Generative Model for Functions
- Reinforcement Learning-Based Coverage Path Planning with Implicit Cellular Decomposition
- Privileged Information Dropout in Reinforcement Learning
- Implicit Policy for Reinforcement Learning
- Reinforcement Learning Based Safe Decision Making for Highway Autonomous Driving
- Locally Persistent Exploration in Continuous Control Tasks with Sparse Rewards
- Hindsight Expectation Maximization for Goal-conditioned Reinforcement Learning
- Distributional Reinforcement Learning with Unconstrained Monotonic Neural Networks
- Challenges of Context and Time in Reinforcement Learning: Introducing Space Fortress as a Benchmark
- Non-local Policy Optimization via Diversity-regularized Collaborative Exploration
- A hierarchical spatial-aware algorithm with efficient reinforcement learning for human-robot task planning and allocation in production
- Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads
- Efficient Model-Free Reinforcement Learning Using Gaussian Process
- Self-organization of action hierarchy and compositionality by reinforcement learning with recurrent neural networks
- CROP: Certifying Robust Policies for Reinforcement Learning through Functional Smoothing
- Variance Reduction for Deep Q-Learning using Stochastic Recursive Gradient
- Learning to Score Behaviors for Guided Policy Optimization
- On the Sample Complexity of Reinforcement Learning with Policy Space Generalization
- Implicit Generative Modeling for Efficient Exploration
- Correlation-aware Cooperative Multigroup Broadcast 360° Video Delivery Network: A Hierarchical Deep Reinforcement Learning Approach
- Finding Needles in a Moving Haystack: Prioritizing Alerts with Adversarial Reinforcement Learning
- Improving the Diversity of Bootstrapped DQN by Replacing Priors With Noise
- The Monte Carlo Transformer: a stochastic self-attention model for sequence prediction
- How Transferable are the Representations Learned by Deep Q Agents?
- Auto-Agent-Distiller: Towards Efficient Deep Reinforcement Learning Agents via Neural Architecture Search
- Empowerment-driven Exploration using Mutual Information Estimation
- Simulating multi-exit evacuation using deep reinforcement learning
- Learning to Represent Action Values as a Hypergraph on the Action Vertices
- Learning Diverse Policies with Soft Self-Generated Guidance
- Explainable Deep Reinforcement Learning Using Introspection in a Non-episodic Task
- Reinforcement Learning with Latent Flow
- Provably Efficient Exploration for Reinforcement Learning Using Unsupervised Learning
- Temporally-Extended ε-Greedy Exploration
- Approximating two value functions instead of one: towards characterizing a new family of Deep Reinforcement Learning algorithms
- Switching Isotropic and Directional Exploration with Parameter Space Noise in Deep Reinforcement Learning
- Accelerating Reinforcement Learning with a Directional-Gaussian-Smoothing Evolution Strategy
- Fractal AI: A fragile theory of intelligence
- Variational Bayes: A report on approaches and applications
- MDP Playground: An Analysis and Debug Testbed for Reinforcement Learning
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- Transfer Heterogeneous Knowledge Among Peer-to-Peer Teammates: A Model Distillation Approach
- Adversary Agnostic Robust Deep Reinforcement Learning
- Improving width-based planning with compact policies
- Continuous Control With Ensemble Deep Deterministic Policy Gradients
- Genetic-Gated Networks for Deep Reinforcement
- Evolutionary Self-Replication as a Mechanism for Producing Artificial Intelligence
- Off-policy Reinforcement Learning with Optimistic Exploration and Distribution Correction
- ExTra: Transfer-guided Exploration
- Improving Experience Replay through Modeling of Similar Transitions' Sets
- Interactive Machine Comprehension with Information Seeking Agents
- Modelling resource allocation in uncertain system environment through deep reinforcement learning
- Solving Atari Games Using Fractals And Entropy
- Thompson Sampling via Local Uncertainty
- Decentralized Deep Reinforcement Learning for Network Level Traffic Signal Control
- Efficient Reinforcement Learning Development with RLzoo
- Lineage Evolution Reinforcement Learning
- Count-Based Temperature Scheduling for Maximum Entropy Reinforcement Learning
- Learning Efficient and Effective Exploration Policies with Counterfactual Meta Policy
- Large-Scale Multi-Agent Deep FBSDEs
- Parameterized Indexed Value Function for Efficient Exploration in Reinforcement Learning
- Coagent Networks Revisited
- CoachNet: An Adversarial Sampling Approach for Reinforcement Learning
- Utilizing Skipped Frames in Action Repeats via Pseudo-Actions
- High Performance Across Two Atari Paddle Games Using the Same Perceptual Control Architecture Without Training
- Transferring Deep Reinforcement Learning with Adversarial Objective and Augmentation
- Learning Transferable Concepts in Deep Reinforcement Learning
- Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning
- Cautious Policy Programming: Exploiting KL Regularization in Monotonic Policy Improvement for Reinforcement Learning