High-Dimensional Continuous Control Using Generalized Advantage Estimation
arXiv:1506.02438
Abstract
Policy gradient methods are an appealing approach in reinforcement learning because they directly optimize the cumulative reward and can straightforwardly be used with nonlinear function approximators such as neural networks. The two main challenges are the large number of samples typically required, and the difficulty of obtaining stable and steady improvement despite the nonstationarity of the incoming data. We address the first challenge by using value functions to substantially reduce the variance of policy gradient estimates at the cost of some bias, with an exponentially-weighted estimator of the advantage function that is analogous to TD(lambda). We address the second challenge by using trust region optimization procedure for both the policy and the value function, which are represented by neural networks. Our approach yields strong empirical results on highly challenging 3D locomotion tasks, learning running gaits for bipedal and quadrupedal simulated robots, and learning a policy for getting the biped to stand up from starting out lying on the ground. In contrast to a body of prior work that uses hand-crafted policy representations, our neural network policies map directly from raw kinematics to joint torques. Our algorithm is fully model-free, and the amount of simulated experience required for the learning tasks on 3D bipeds corresponds to 1-2 weeks of real time.
References in corpus (3)
Cited by in corpus (392)
- A Brief Survey of Deep Reinforcement Learning
- Asynchronous Methods for Deep Reinforcement Learning
- Learning agile and dynamic motor skills for legged robots
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Benchmarking Deep Reinforcement Learning for Continuous Control
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
- Emergence of Locomotion Behaviours in Rich Environments
- Fast Adaptive Task Offloading in Edge Computing based on Meta Reinforcement Learning
- A Closer Look at Invalid Action Masking in Policy Gradient Algorithms
- AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control
- Emergent Tool Use From Multi-Agent Autocurricula
- Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
- Generative Adversarial Imitation Learning
- Fully Decentralized Multi-Agent Reinforcement Learning with Networked Agents
- Alleviating catastrophic forgetting using context-dependent gating and synaptic stabilization
- SNAS: Stochastic Neural Architecture Search
- Learning Dexterous In-Hand Manipulation
- Reward Constrained Policy Optimization
- A Survey of Deep RL and IL for Autonomous Driving Policy Learning
- Sample Efficient Actor-Critic with Experience Replay
- Flow: A Modular Learning Framework for Mixed Autonomy Traffic
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Episodic Curiosity through Reachability
- Equivalence Between Policy Gradients and Soft Q-Learning
- Simple random search provides a competitive approach to reinforcement learning
- Mobile Augmented Reality: User Interfaces, Frameworks, and Intelligence
- Lyapunov-based Safe Policy Optimization for Continuous Control
- Continuous Adaptation via Meta-Learning in Nonstationary and Competitive Environments
- Learning human behaviors from motion capture by adversarial imitation
- Third-Person Imitation Learning
- Emergent Complexity via Multi-Agent Competition
- Automatic Goal Generation for Reinforcement Learning Agents
- Unsupervised Predictive Memory in a Goal-Directed Agent
- Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
- InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations
- Reverse Curriculum Generation for Reinforcement Learning
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- Generative Design by Reinforcement Learning: Enhancing the Diversity of Topology Optimization Designs
- Neural SLAM: Learning to Explore with External Memory
- Tianshou: a Highly Modularized Deep Reinforcement Learning Library
- What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
- RUDDER: Return Decomposition for Delayed Rewards
- Learning and Transfer of Modulated Locomotor Controllers
- Deep Attention Recurrent Q-Network
- Feature Control as Intrinsic Motivation for Hierarchical Reinforcement Learning
- Backpropagation through the Void: Optimizing control variates for black-box gradient estimation
- Efficient Exploration via State Marginal Matching
- Controlling an Autonomous Vehicle with Deep Reinforcement Learning
- HAIM-DRL: Enhanced Human-in-the-loop Reinforcement Learning for Safe and Efficient Autonomous Driving
- Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
- Deep Reinforcement Learning for Trajectory Path Planning and Distributed Inference in Resource-Constrained UAV Swarms
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- Robust High-speed Running for Quadruped Robots via Deep Reinforcement Learning
- Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning
- Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control
- Online Service Migration in Mobile Edge with Incomplete System Information: A Deep Recurrent Actor-Critic Learning Approach
- Queueing Network Controls via Deep Reinforcement Learning
- Deep Reinforcement Learning for Event-Driven Multi-Agent Decision Processes
- Improving Generalization in Meta Reinforcement Learning using Learned Objectives
- Unified Automatic Control of Vehicular Systems with Reinforcement Learning
- Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning
- TStarBots: Defeating the Cheating Level Builtin AI in StarCraft II in the Full Game
- Learning Algorithms for Active Learning
- Parameter Sharing Deep Deterministic Policy Gradient for Cooperative Multi-agent Reinforcement Learning
- Deep Reinforcement Learning Algorithm for Dynamic Pricing of Express Lanes with Multiple Access Locations
- A differentiable programming method for quantum control
- Reinforcement Learning for Generative AI: State of the Art, Opportunities and Open Research Challenges
- Phasic Policy Gradient
- Dealing with Sparse Rewards in Reinforcement Learning
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- Deep Reinforcement Learning for Cybersecurity Threat Detection and Protection: A Review
- Improving Generalization in Reinforcement Learning with Mixture Regularization
- Boosting 5G on Smart Grid Communication: A Smart RAN Slicing Approach
- Almost Optimal Model-Free Reinforcement Learning via Reference-Advantage Decomposition
- Understanding the impact of entropy on policy optimization
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- Preparing for the Unknown: Learning a Universal Policy with Online System Identification
- Diversity Policy Gradient for Sample Efficient Quality-Diversity Optimization
- Resource Optimization for Semantic-Aware Networks with Task Offloading
- Search on the Replay Buffer: Bridging Planning and Reinforcement Learning
- Constrained Attractor Selection Using Deep Reinforcement Learning
- Deep Implicit Coordination Graphs for Multi-agent Reinforcement Learning
- Sample Efficient Policy Gradient Methods with Recursive Variance Reduction
- Enforcing Policy Feasibility Constraints through Differentiable Projection for Energy Optimization
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames
- Active Neural Localization
- Towards Optimally Decentralized Multi-Robot Collision Avoidance via Deep Reinforcement Learning
- Residual Force Control for Agile Human Behavior Imitation and Extended Motion Synthesis
- ENERO: Efficient Real-Time WAN Routing Optimization with Deep Reinforcement Learning
- TD-Regularized Actor-Critic Methods
- Recurrent Deterministic Policy Gradient Method for Bipedal Locomotion on Rough Terrain Challenge
- Conservative Safety Critics for Exploration
- Redirection Controller Using Reinforcement Learning
- One model Packs Thousands of Items with Recurrent Conditional Query Learning
- Deep Reinforcement Learning for Six Degree-of-Freedom Planetary Powered Descent and Landing
- Zero Time Waste: Recycling Predictions in Early Exit Neural Networks
- Edge Generation Scheduling for DAG Tasks Using Deep Reinforcement Learning
- Trajectory Planning with Deep Reinforcement Learning in High-Level Action Spaces
- On Learning Intrinsic Rewards for Policy Gradient Methods
- SCC: an efficient deep reinforcement learning agent mastering the game of StarCraft II
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers
- A Deep Multi-Agent Reinforcement Learning Approach to Autonomous Separation Assurance
- Stochastic optimal well control in subsurface reservoirs using reinforcement learning
- Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning
- A Survey of Embodied AI: From Simulators to Research Tasks
- An Adaptive Clipping Approach for Proximal Policy Optimization
- Mastering Complex Control in MOBA Games with Deep Reinforcement Learning
- Enhancing Efficiency and Propulsion in Bio-mimetic Robotic Fish through End-to-End Deep Reinforcement Learning
- TRC: Trust Region Conditional Value at Risk for Safe Reinforcement Learning
- Settling the Variance of Multi-Agent Policy Gradients
- Learning Event-triggered Control from Data through Joint Optimization
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- A Cooperative-Competitive Multi-Agent Framework for Auto-bidding in Online Advertising
- Generalized Planning With Deep Reinforcement Learning
- Mastering Atari with Discrete World Models
- Cooperative Assistance in Robotic Surgery through Multi-Agent Reinforcement Learning
- Benchmarking Potential Based Rewards for Learning Humanoid Locomotion
- Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives
- Learning Intrusion Prevention Policies through Optimal Stopping
- Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
- Meta-Learning through Hebbian Plasticity in Random Networks
- A Tour of Reinforcement Learning: The View from Continuous Control
- Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation
- Optimistic Reinforcement Learning by Forward Kullback-Leibler Divergence Optimization
- Generalization in Transfer Learning
- The Faults in Our Pi Stars: Security Issues and Open Challenges in Deep Reinforcement Learning
- Robust Model-based Reinforcement Learning for Autonomous Greenhouse Control
- Tsallis Reinforcement Learning: A Unified Framework for Maximum Entropy Reinforcement Learning
- GOPT: Generalizable Online 3D Bin Packing via Transformer-based Deep Reinforcement Learning
- Beyond the One Step Greedy Approach in Reinforcement Learning
- What Matters for Adversarial Imitation Learning?
- Iroko: A Framework to Prototype Reinforcement Learning for Data Center Traffic Control
- Collaborative Deep Reinforcement Learning
- Collaborative Visual Navigation
- A Survey on Autonomous Vehicle Control in the Era of Mixed-Autonomy: From Physics-Based to AI-Guided Driving Policy Learning
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- Towards Cooperation in Sequential Prisoner's Dilemmas: a Deep Multiagent Reinforcement Learning Approach
- Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods
- Scalable Centralized Deep Multi-Agent Reinforcement Learning via Policy Gradients
- Modern Deep Reinforcement Learning Algorithms
- Deep Reinforcement Learning for Dexterous Manipulation with Concept Networks
- Barrier Functions Inspired Reward Shaping for Reinforcement Learning
- Obtaining Robust Control and Navigation Policies for Multi-Robot Navigation via Deep Reinforcement Learning
- Bayesian Policy Gradients via Alpha Divergence Dropout Inference
- Model-based Reinforcement Learning from Signal Temporal Logic Specifications
- Neural Logic Reinforcement Learning
- Adaptive Guidance with Reinforcement Meta-Learning
- Single-Timescale Stochastic Nonconvex-Concave Optimization for Smooth Nonlinear TD Learning
- Human-Inspired Multi-Agent Navigation using Knowledge Distillation
- Self-Attentional Credit Assignment for Transfer in Reinforcement Learning
- Falsification-Based Robust Adversarial Reinforcement Learning
- RL-LABEL: A Deep Reinforcement Learning Approach Intended for AR Label Placement in Dynamic Scenarios
- Learning Belief Representations for Imitation Learning in POMDPs
- Applications of Multi-Agent Reinforcement Learning in Future Internet: A Comprehensive Survey
- Learning Heuristic Search via Imitation
- Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph Learning
- Attractor Selection in Nonlinear Energy Harvesting Using Deep Reinforcement Learning
- Intervention Aided Reinforcement Learning for Safe and Practical Policy Optimization in Navigation
- Learning to Navigate Using Mid-Level Visual Priors
- Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
- -GAIL: Learning -Divergence for Generative Adversarial Imitation Learning
- Using Simulation Optimization to Improve Zero-shot Policy Transfer of Quadrotors
- A Deep Reinforcement Learning-based Adaptive Charging Policy for WRSNs
- A Reinforced Generation of Adversarial Examples for Neural Machine Translation
- SpikePropamine: Differentiable Plasticity in Spiking Neural Networks
- A Safe Reinforcement Learning Algorithm for Supervisory Control of Power Plants
- Actor-Critic Reinforcement Learning for Control with Stability Guarantee
- PyTester: Deep Reinforcement Learning for Text-to-Testcase Generation
- End-to-End Reinforcement Learning of Koopman Models for Economic Nonlinear Model Predictive Control
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian Optimization
- A Deep Reinforcement Learning Framework for Eco-driving in Connected and Automated Hybrid Electric Vehicles
- Effective Scheduling Function Design in SDN through Deep Reinforcement Learning
- Robust Policy Gradient against Strong Data Corruption
- Cautious Reinforcement Learning via Distributional Risk in the Dual Domain
- MAT: Multi-Fingered Adaptive Tactile Grasping via Deep Reinforcement Learning
- Brick-by-Brick: Combinatorial Construction with Deep Reinforcement Learning
- Traffic Signal Cycle Control with Centralized Critic and Decentralized Actors under Varying Intervention Frequencies
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft
- Smooth Exploration for Robotic Reinforcement Learning
- Online Prediction-Assisted Safe Reinforcement Learning for Electric Vehicle Charging Station Recommendation in Dynamically Coupled Transportation-Power Systems
- Quadratic Q-network for Learning Continuous Control for Autonomous Vehicles
- Deep Reinforcement Learning Controller for 3D Path-following and Collision Avoidance by Autonomous Underwater Vehicles
- Multi-Agent Deep Reinforcement Learning for Request Dispatching in Distributed-Controller Software-Defined Networking
- Proximal Policy Optimization with Adaptive Threshold for Symmetric Relative Density Ratio
- An Empirical Study on Google Research Football Multi-agent Scenarios
- Deep Reinforcement Learning to Acquire Navigation Skills for Wheel-Legged Robots in Complex Environments
- Derivative-Free Policy Optimization for Linear Risk-Sensitive and Robust Control Design: Implicit Regularization and Sample Complexity
- Direct and indirect reinforcement learning
- Meta-AAD: Active Anomaly Detection with Deep Reinforcement Learning
- Adversarial Reinforcement Learning for Observer Design in Autonomous Systems under Cyber Attacks
- TiKick: Towards Playing Multi-agent Football Full Games from Single-agent Demonstrations
- Learning Symbolic Rules for Interpretable Deep Reinforcement Learning
- Hindsight Trust Region Policy Optimization
- OffWorld Gym: open-access physical robotics environment for real-world reinforcement learning benchmark and research
- Value Iteration in Continuous Actions, States and Time
- Discovering Command and Control Channels Using Reinforcement Learning
- Safe Reinforcement Learning with Natural Language Constraints
- Marginalized State Distribution Entropy Regularization in Policy Optimization
- Deep Reinforcement Learning for Multi-Driver Vehicle Dispatching and Repositioning Problem
- Entanglement engineering of optomechanical systems by reinforcement learning
- Combined Peak Reduction and Self-Consumption Using Proximal Policy Optimization
- Policy Search with Rare Significant Events: Choosing the Right Partner to Cooperate with
- Emergent Road Rules In Multi-Agent Driving Environments
- Augmenting GAIL with BC for sample efficient imitation learning
- Reinforcement Learning with Function Approximation: From Linear to Nonlinear
- Privacy-preserving Q-Learning with Functional Noise in Continuous State Spaces
- Application of Self-Play Reinforcement Learning to a Four-Player Game of Imperfect Information
- Expert-augmented actor-critic for ViZDoom and Montezumas Revenge
- Reinforcement Learning with Adaptive Curriculum Dynamics Randomization for Fault-Tolerant Robot Control
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement Learning
- Learning Cooperative Multi-Agent Policies with Partial Reward Decoupling
- Unsupervised Domain Adaptation for Visual Navigation
- Cascade Attribute Learning Network
- Decentralized Reinforcement Learning: Global Decision-Making via Local Economic Transactions
- Run, skeleton, run: skeletal model in a physics-based simulation
- Robust Asymmetric Learning in POMDPs
- Non-local Policy Optimization via Diversity-regularized Collaborative Exploration
- Learning Policies from Self-Play with Policy Gradients and MCTS Value Estimates
- Interactive Visualization for Debugging RL
- Adversarial joint attacks on legged robots
- Hierarchical principles of embodied reinforcement learning: A review
- Variational Autoencoders for Opponent Modeling in Multi-Agent Systems
- Energy Minimization in UAV-Aided Networks: Actor-Critic Learning for Constrained Scheduling Optimization
- Is the Policy Gradient a Gradient?
- A Deeper Look at Discounting Mismatch in Actor-Critic Algorithms
- Action Guidance: Getting the Best of Sparse Rewards and Shaped Rewards for Real-time Strategy Games
- Mutual Information Based Knowledge Transfer Under State-Action Dimension Mismatch
- Trust Region Value Optimization using Kalman Filtering
- MRAC-RL: A Framework for On-Line Policy Adaptation Under Parametric Model Uncertainty
- KoGuN: Accelerating Deep Reinforcement Learning via Integrating Human Suboptimal Knowledge
- Only Relevant Information Matters: Filtering Out Noisy Samples to Boost RL
- Challenges of Context and Time in Reinforcement Learning: Introducing Space Fortress as a Benchmark
- Prioritized Level Replay
- Proportional integral derivative controller assisted reinforcement learning for path following by autonomous underwater vehicles
- Preventing Imitation Learning with Adversarial Policy Ensembles
- Adaptive Stress Testing for Autonomous Vehicles
- On-Policy Trust Region Policy Optimisation with Replay Buffers
- Deep reinforcement learning for feedback control in a collective flashing ratchet
- AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control
- Loaded DiCE: Trading off Bias and Variance in Any-Order Score Function Estimators for Reinforcement Learning
- A Risk-Sensitive Approach to Policy Optimization
- GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model
- Stochastic Particle-Optimization Sampling and the Non-Asymptotic Convergence Theory
- Agent Modelling under Partial Observability for Deep Reinforcement Learning
- GRN: Generative Rerank Network for Context-wise Recommendation
- P3O: Policy-on Policy-off Policy Optimization
- REPAINT: Knowledge Transfer in Deep Reinforcement Learning
- Learning Value Functions in Deep Policy Gradients using Residual Variance
- Backprop-Q: Generalized Backpropagation for Stochastic Computation Graphs
- Mean-Semivariance Policy Optimization via Risk-Averse Reinforcement Learning
- Learning to Navigate Cloth using Haptics
- Reinforcement Re-ranking with 2D Grid-based Recommendation Panels
- Learning Powerful Policies by Using Consistent Dynamics Model
- Relationship Explainable Multi-objective Reinforcement Learning with Semantic Explainability Generation
- Role Playing Learning for Socially Concomitant Mobile Robot Navigation
- RMP2: A Structured Composable Policy Class for Robot Learning
- Bridging the Imitation Gap by Adaptive Insubordination
- Augmented Random Search for Quadcopter Control: An alternative to Reinforcement Learning
- Harnessing Distribution Ratio Estimators for Learning Agents with Quality and Diversity
- Towards Automatic Actor-Critic Solutions to Continuous Control
- Hierarchical Expert Networks for Meta-Learning
- Meta Reinforcement Learning with Autonomous Inference of Subtask Dependencies
- Proximal Policy Optimization Smoothed Algorithm
- Policy Search by Target Distribution Learning for Continuous Control
- Reinforcement Learning Meets Hybrid Zero Dynamics: A Case Study for RABBIT
- On Effective Scheduling of Model-based Reinforcement Learning
- Multimodal Reward Shaping for Efficient Exploration in Reinforcement Learning
- Proximal Policy Gradient: PPO with Policy Gradient
- CATCH: Context-based Meta Reinforcement Learning for Transferrable Architecture Search
- Where to go next: Learning a Subgoal Recommendation Policy for Navigation Among Pedestrians
- SimPoE: Simulated Character Control for 3D Human Pose Estimation
- Approximate Newton policy gradient algorithms
- Learning State Representations via Retracing in Reinforcement Learning
- Gym-RTS: Toward Affordable Full Game Real-time Strategy Games Research with Deep Reinforcement Learning
- Improving Learning from Demonstrations by Learning from Experience
- On Proximal Policy Optimization's Heavy-tailed Gradients
- DeepConfig: Automating Data Center Network Topologies Management with Machine Learning
- Dual policy as self-model for planning
- Optimization of a Triangular Delaunay Mesh Generator using Reinforcement Learning
- Amoeba: Circumventing ML-supported Network Censorship via Adversarial Reinforcement Learning
- SeRO: Self-Supervised Reinforcement Learning for Recovery from Out-of-Distribution Situations
- Improving Long-Term Metrics in Recommendation Systems using Short-Horizon Reinforcement Learning
- Hyperparameter optimization with REINFORCE and Transformers
- Hindsight Value Function for Variance Reduction in Stochastic Dynamic Environment
- Expert Level control of Ramp Metering based on Multi-task Deep Reinforcement Learning
- Policy Optimization Reinforcement Learning with Entropy Regularization
- Merging Deterministic Policy Gradient Estimations with Varied Bias-Variance Tradeoff for Effective Deep Reinforcement Learning
- A Human-Centered Data-Driven Planner-Actor-Critic Architecture via Logic Programming
- Smooth Imitation Learning via Smooth Costs and Smooth Policies
- Generative Inverse Deep Reinforcement Learning for Online Recommendation
- Dynamic allocation of limited memory resources in reinforcement learning
- Walking with MIND: Mental Imagery eNhanceD Embodied QA
- Sparse Attention Guided Dynamic Value Estimation for Single-Task Multi-Scene Reinforcement Learning
- A Distance-based Anomaly Detection Framework for Deep Reinforcement Learning
- Dynamic Value Estimation for Single-Task Multi-Scene Reinforcement Learning
- Learning Adaptive Display Exposure for Real-Time Advertising
- NavigationNet: A Large-scale Interactive Indoor Navigation Dataset
- Attention Based Natural Language Grounding by Navigating Virtual Environment
- Competitiveness of MAP-Elites against Proximal Policy Optimization on locomotion tasks in deterministic simulations
- Compatible features for Monotonic Policy Improvement
- Taming an autonomous surface vehicle for path following and collision avoidance using deep reinforcement learning
- Monte-Carlo Tree Search for Policy Optimization
- Recurrent Value Functions
- Direct Advantage Estimation
- Cascade Attribute Network: Decomposing Reinforcement Learning Control Policies using Hierarchical Neural Networks
- Reinforcement Learning for Nested Polar Code Construction
- Decentralized Multi-Agents by Imitation of a Centralized Controller
- An Active Learning Framework for Efficient Robust Policy Search
- An Accelerated Fitted Value Iteration Algorithm for MDPs with Finite and Vector-Valued Action Space
- PFPN: Continuous Control of Physically Simulated Characters using Particle Filtering Policy Network
- Analyzing Visual Representations in Embodied Navigation Tasks
- Smart Scheduling based on Deep Reinforcement Learning for Cellular Networks
- Safety Enhancement for Deep Reinforcement Learning in Autonomous Separation Assurance
- Optimistic Proximal Policy Optimization
- Recomposing the Reinforcement Learning Building Blocks with Hypernetworks
- Controlling an Inverted Pendulum with Policy Gradient Methods-A Tutorial
- Adaptive Coordination Offsets for Signalized Arterial Intersections using Deep Reinforcement Learning
- Deep-Reinforcement-Learning for Gliding and Perching Bodies
- Reparameterized Variational Divergence Minimization for Stable Imitation
- Improving the Generalization of Unseen Crowd Behaviors for Reinforcement Learning based Local Motion Planners
- Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning
- Battlesnake Challenge: A Multi-agent Reinforcement Learning Playground with Human-in-the-loop
- Policy Optimization with Second-Order Advantage Information
- Multi-agent Policy Optimization with Approximatively Synchronous Advantage Estimation
- Configuration Path Control
- Learning Human Behaviors for Robot-Assisted Dressing
- Meta Arcade: A Configurable Environment Suite for Meta-Learning
- Creativity in Robot Manipulation with Deep Reinforcement Learning
- Bregman Gradient Policy Optimization
- Buffer-aware Wireless Scheduling based on Deep Reinforcement Learning
- On the Sample Complexity and Metastability of Heavy-tailed Policy Search in Continuous Control
- Quasi-Newton Trust Region Policy Optimization
- Learning to drive via Apprenticeship Learning and Deep Reinforcement Learning
- IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target Networks
- Expanding Motor Skills through Relay Neural Networks
- Discrete Action On-Policy Learning with Action-Value Critic
- Deep Reinforcement Learning for URLLC data management on top of scheduled eMBB traffic
- Coordinated Proximal Policy Optimization
- Design of AoI-Aware 5G Uplink Scheduler UsingReinforcement Learning
- Learning to Mix n-Step Returns: Generalizing lambda-Returns for Deep Reinforcement Learning
- Variance Reduced Advantage Estimation with Hindsight Credit Assignment
- Improving the sample-efficiency of neural architecture search with reinforcement learning
- Learning Policies through Quantile Regression
- Zero-shot Policy Learning with Spatial Temporal RewardDecomposition on Contingency-aware Observation
- Cooperative and Asynchronous Transformer-based Mission Planning for Heterogeneous Teams of Mobile Robots
- Deep RL Agent for a Real-Time Action Strategy Game
- Enhanced Scene Specificity with Sparse Dynamic Value Estimation
- Enhancing Reinforcement Learning Through Guided Search
- Learning Time-Sensitive Strategies in Space Fortress
- A Joint Planning and Learning Framework for Human-Aided Decision-Making
- Model-free Policy Learning with Reward Gradients
- Learning to Reach, Swim, Walk and Fly in One Trial: Data-Driven Control with Scarce Data and Side Information
- Predictive Synthesis of Quantum Materials by Probabilistic Reinforcement Learning
- Emergence of Different Modes of Tool Use in a Reaching and Dragging Task
- Hindsight Reward Tweaking via Conditional Deep Reinforcement Learning
- MBDP: A Model-based Approach to Achieve both Robustness and Sample Efficiency via Double Dropout Planning
- CubeTR: Learning to Solve The Rubiks Cube Using Transformers
- Graph Convolutional Policy for Solving Tree Decomposition via Reinforcement Learning Heuristics
- Coordinate-wise Control Variates for Deep Policy Gradients
- Learning to Manipulate Amorphous Materials
- Sample Complexity of Estimating the Policy Gradient for Nearly Deterministic Dynamical Systems
- Variational Inference for Policy Gradient
- On The Transferability of Deep-Q Networks
- Easy Monotonic Policy Iteration
- Biased Estimates of Advantages over Path Ensembles
- A2: Extracting Cyclic Switchings from DOB-nets for Rejecting Excessive Disturbances
- Self-Constructing Neural Networks Through Random Mutation
- Improving Automatic Source Code Summarization via Deep Reinforcement Learning
- A Validated Physical Model For Real-Time Simulation of Soft Robotic Snakes
- Adaptation of Quadruped Robot Locomotion with Meta-Learning
- Relabel the Noise: Joint Extraction of Entities and Relations via Cooperative Multiagents
- Out-of-the-box channel pruned networks
- Efficiently Training On-Policy Actor-Critic Networks in Robotic Deep Reinforcement Learning with Demonstration-like Sampled Exploration
- Batch-Augmented Multi-Agent Reinforcement Learning for Efficient Traffic Signal Optimization
- Building Intelligent Autonomous Navigation Agents
- Injective State-Image Mapping facilitates Visual Adversarial Imitation Learning
- Polymatrix Competitive Gradient Descent
- Zero-shot generalization using cascaded system-representations
- Sequential Coordination of Deep Models for Learning Visual Arithmetic
- Imitation Learning via Simultaneous Optimization of Policies and Auxiliary Trajectories
- Explaining Fast Improvement in Online Imitation Learning
- Lifetime policy reuse and the importance of task capacity
- Online Algorithms and Policies Using Adaptive and Machine Learning Approaches
- Learning in Sparse Rewards settings through Quality-Diversity algorithms
- EnTRPO: Trust Region Policy Optimization Method with Entropy Regularization
- Model-Free Synthesis via Adversarial Reinforcement Learning
- GO Hessian for Expectation-Based Objectives
- Discrete linear-complexity reinforcement learning in continuous action spaces for Q-learning algorithms
- Crowdfunding Dynamics Tracking: A Reinforcement Learning Approach
- Dynamic Matching Markets in Power Grid: Concepts and Solution using Deep Reinforcement Learning
- Spatially and Seamlessly Hierarchical Reinforcement Learning for State Space and Policy space in Autonomous Driving