Continuous control with deep reinforcement learning
arXiv:1509.02971
Abstract
We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algorithm robustly solves more than 20 simulated physics tasks, including classic problems such as cartpole swing-up, dexterous manipulation, legged locomotion and car driving. Our algorithm is able to find policies whose performance is competitive with those found by a planning algorithm with full access to the dynamics of the domain and its derivatives. We further demonstrate that for many of the tasks the algorithm can learn policies end-to-end: directly from raw pixel inputs.
10 pages + supplementary
References in corpus (6)
- Adam: A Method for Stochastic Optimization
- End-to-End Training of Deep Visuomotor Policies
- Learning Continuous Control Policies by Stochastic Value Gradients
- Memory-based control with recurrent neural networks
- Gradient Estimation Using Stochastic Computation Graphs
- Compatible Value Gradients for Reinforcement Learning of Continuous Deep Policies
Cited by in corpus (1632)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Bootstrap your own latent: A new approach to self-supervised Learning
- Addressing Function Approximation Error in Actor-Critic Methods
- Soft Actor-Critic Algorithms and Applications
- A Survey of Deep Learning Techniques for Autonomous Driving
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Asynchronous Methods for Deep Reinforcement Learning
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- End-to-End Training of Deep Visuomotor Policies
- Deep Reinforcement Learning for Multi-Agent Systems: A Review of Challenges, Solutions and Applications
- AMC: AutoML for Model Compression and Acceleration on Mobile Devices
- Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Applications of Deep Learning and Reinforcement Learning to Biological Data
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
- Human-level performance in first-person multiplayer games with population-based deep reinforcement learning
- A Convergence Theory for Deep Learning via Over-Parameterization
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Emergence of Locomotion Behaviours in Rich Environments
- Deep Reinforcement Learning for Cyber Security
- Artificial Neural Networks trained through Deep Reinforcement Learning discover control strategies for active flow control
- Noisy Networks for Exploration
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- The Marginal Value of Adaptive Gradient Methods in Machine Learning
- Deep Reinforcement Learning: An Overview
- DeepMind Control Suite
- Exploration in Deep Reinforcement Learning: A Survey
- Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
- Robust Physical-World Attacks on Deep Learning Models
- A Review of Deep Reinforcement Learning for Smart Building Energy Management
- Deep Reinforcement Learning for Smart Home Energy Management
- Safe, Efficient, and Comfortable Velocity Control based on Reinforcement Learning for Autonomous Driving
- MolGAN: An implicit generative model for small molecular graphs
- Deep Reinforcement Learning for Page-wise Recommendations
- Fast Adaptive Task Offloading in Edge Computing based on Meta Reinforcement Learning
- Reinforcement Learning with Deep Energy-Based Policies
- Robust Adversarial Reinforcement Learning
- How Generative Adversarial Networks and Their Variants Work: An Overview
- Learning Continuous Control Policies by Stochastic Value Gradients
- Hindsight Experience Replay
- Distributional Soft Actor-Critic: Off-Policy Reinforcement Learning for Addressing Value Estimation Errors
- Fully Decentralized Multi-Agent Reinforcement Learning with Networked Agents
- #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
- Distributed Distributional Deterministic Policy Gradients
- Continuous Deep Q-Learning with Model-based Acceleration
- Combining Planning and Deep Reinforcement Learning in Tactical Decision Making for Autonomous Driving
- GCN-RL Circuit Designer: Transferable Transistor Sizing with Graph Neural Networks and Reinforcement Learning
- State Representation Learning for Control: An Overview
- Learning Attentional Communication for Multi-Agent Cooperation
- A Survey of End-to-End Driving: Architectures and Training Methods
- Actor-Attention-Critic for Multi-Agent Reinforcement Learning
- Safe Exploration in Continuous Action Spaces
- Off-Policy Deep Reinforcement Learning without Exploration
- Graph networks as learnable physics engines for inference and control
- Optimal and Scalable Caching for 5G Using Reinforcement Learning of Space-time Popularities
- Multiagent Bidirectionally-Coordinated Nets: Emergence of Human-level Coordination in Learning to Play StarCraft Combat Games
- Challenges of Real-World Reinforcement Learning
- Behavior Regularized Offline Reinforcement Learning
- MOPO: Model-based Offline Policy Optimization
- Reinforcement Learning with Augmented Data
- Benchmarking Model-Based Reinforcement Learning
- A Survey of Deep RL and IL for Autonomous Driving Policy Learning
- Comparison of Deep Reinforcement Learning and Model Predictive Control for Adaptive Cruise Control
- Uncertainty-Aware Reinforcement Learning for Collision Avoidance
- One-Shot Imitation Learning
- An Actor-Critic Algorithm for Sequence Prediction
- Deep Reinforcement Learning Attitude Control of Fixed-Wing UAVs Using Proximal Policy Optimization
- Decentralized Computation Offloading for Multi-User Mobile Edge Computing: A Deep Reinforcement Learning Approach
- Reinforcement Learning for IoT Security: A Comprehensive Survey
- CURL: Contrastive Unsupervised Representations for Reinforcement Learning
- Federated Reinforcement Learning: Techniques, Applications, and Open Challenges
- A Survey of Machine Learning for Computer Architecture and Systems
- Memory-based control with recurrent neural networks
- Sample Efficient Actor-Critic with Experience Replay
- Go-Explore: a New Approach for Hard-Exploration Problems
- Hierarchical Reinforcement Learning for Self-Driving Decision-Making without Reliance on Labeled Driving Data
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
- Differentiable MPC for End-to-end Planning and Control
- dm_control: Software and Tasks for Continuous Control
- Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning
- Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
- Resource Management in Wireless Networks via Multi-Agent Deep Reinforcement Learning
- Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
- Large-Scale Traffic Signal Control Using a Novel Multi-Agent Reinforcement Learning
- Stochastic Neural Networks for Hierarchical Reinforcement Learning
- Sim4CV: A Photo-Realistic Simulator for Computer Vision Applications
- A Deeper Look at Experience Replay
- Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Ablation Studies in Artificial Neural Networks
- Meta-Reinforcement Learning of Structured Exploration Strategies
- A Review of Tracking, Prediction and Decision Making Methods for Autonomous Driving
- A Deep-Reinforcement Learning Approach for Software-Defined Networking Routing Optimization
- Real-Time Bidding with Multi-Agent Reinforcement Learning in Display Advertising
- A Minimalist Approach to Offline Reinforcement Learning
- Control of Memory, Active Perception, and Action in Minecraft
- Simple random search provides a competitive approach to reinforcement learning
- Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Recent advances in applying deep reinforcement learning for flow control: perspectives and future directions
- MOReL : Model-Based Offline Reinforcement Learning
- Learning Force Control for Contact-rich Manipulation Tasks with Rigid Position-controlled Robots
- Survey on reinforcement learning for language processing
- Recent Advances and Applications of Machine Learning in Experimental Solid Mechanics: A Review
- Temporal Difference Models: Model-Free Deep RL for Model-Based Control
- Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space
- Transfer Learning in Deep Reinforcement Learning: A Survey
- The Predictron: End-To-End Learning and Planning
- A Survey of Deep Reinforcement Learning in Video Games
- Virtual to Real Reinforcement Learning for Autonomous Driving
- Adaptive Traffic Signal Control: Deep Reinforcement Learning Algorithm with Experience Replay and Target Network
- Optimal control towards sustainable wastewater treatment plants based on multi-agent reinforcement learning
- Perception and Navigation in Autonomous Systems in the Era of Learning: A Survey
- Deep Reinforcement Learning for Process Control: A Primer for Beginners
- Game-Theoretic Multiagent Reinforcement Learning
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- Emergent Complexity via Multi-Agent Competition
- Deep Deterministic Policy Gradient for Urban Traffic Light Control
- Automatic Goal Generation for Reinforcement Learning Agents
- Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
- RLOC: Terrain-Aware Legged Locomotion using Reinforcement Learning and Optimal Control
- Collective Robot Reinforcement Learning with Distributed Asynchronous Guided Policy Search
- InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations
- Symplectic ODE-Net: Learning Hamiltonian Dynamics with Control
- Dream to Control: Learning Behaviors by Latent Imagination
- Reverse Curriculum Generation for Reinforcement Learning
- Optimization for deep learning: theory and algorithms
- Evolved Policy Gradients
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
- When does reinforcement learning stand out in quantum control? A comparative study on state preparation
- Learning to Poke by Poking: Experiential Learning of Intuitive Physics
- Rapid trial-and-error learning with simulation supports flexible tool use and physical reasoning
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- Deep Reinforcement Learning with Shallow Controllers: An Experimental Application to PID Tuning
- Intelligent Inverse Treatment Planning via Deep Reinforcement Learning, a Proof-of-Principle Study in High Dose-rate Brachytherapy for Cervical Cancer
- A Theoretical Analysis of Deep Q-Learning
- Decentralized Power Allocation for MIMO-NOMA Vehicular Edge Computing Based on Deep Reinforcement Learning
- Robust Deep Reinforcement Learning with Adversarial Attacks
- Scalable agent alignment via reward modeling: a research direction
- Reinforcement Learning in Healthcare: A Survey
- When to Trust Your Model: Model-Based Policy Optimization
- A Survey of Algorithms for Black-Box Safety Validation of Cyber-Physical Systems
- Constrained Policy Optimization
- On the Convergence Rate of Training Recurrent Neural Networks
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
- Reinforcement Learning Algorithms: An Overview and Classification
- Deep Reinforcement Learning for List-wise Recommendations
- Review: Deep Learning in Electron Microscopy
- A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation
- Learning Visual Predictive Models of Physics for Playing Billiards
- FACMAC: Factored Multi-Agent Centralised Policy Gradients
- Physical Adversarial Examples for Object Detectors
- Learning and Transfer of Modulated Locomotor Controllers
- Machine Learning for the Control and Monitoring of Electric Machine Drives: Advances and Trends
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning
- Virtual-to-real Deep Reinforcement Learning: Continuous Control of Mobile Robots for Mapless Navigation
- Intelligent Electric Vehicle Charging Recommendation Based on Multi-Agent Reinforcement Learning
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- PrecoderNet: Hybrid Beamforming for Millimeter Wave Systems with Deep Reinforcement Learning
- An Optimistic Perspective on Offline Reinforcement Learning
- Guided Deep Reinforcement Learning for Swarm Systems
- Deep Reinforcement Learning for Black-Box Testing of Android Apps
- Sim-to-Real Reinforcement Learning for Deformable Object Manipulation
- Neural Network-based Flight Control Systems: Present and Future
- Programmatically Interpretable Reinforcement Learning
- CEM-RL: Combining evolutionary and gradient-based methods for policy search
- Connecting Generative Adversarial Networks and Actor-Critic Methods
- Energy Management Based on Multi-Agent Deep Reinforcement Learning for A Multi-Energy Industrial Park
- How to Certify Machine Learning Based Safety-critical Systems? A Systematic Literature Review
- A survey on intrinsic motivation in reinforcement learning
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- Machine and Deep Learning for IoT Security and Privacy: Applications, Challenges, and Future Directions
- Combining policy gradient and Q-learning
- Efficient Exploration via State Marginal Matching
- Global optimization of quantum dynamics with AlphaZero deep exploration
- Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning
- CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning Algorithms
- Controlling an Autonomous Vehicle with Deep Reinforcement Learning
- Horizon: Facebook's Open Source Applied Reinforcement Learning Platform
- Learning Predictive Representations for Deformable Objects Using Contrastive Estimation
- Deep Reinforcement Learning based Recommendation with Explicit User-Item Interactions Modeling
- Monte Carlo Gradient Estimation in Machine Learning
- Deep Graph Convolutional Reinforcement Learning for Financial Portfolio Management -- DeepPocket
- Artificial Intelligence and its Role in Near Future
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- Proximal Distilled Evolutionary Reinforcement Learning
- Towards Cognitive Exploration through Deep Reinforcement Learning for Mobile Robots
- Exploring Model-based Planning with Policy Networks
- GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
- FIRM: An Intelligent Fine-Grained Resource Management Framework for SLO-Oriented Microservices
- Learning Physical Intuition of Block Towers by Example
- Mapless Navigation among Dynamics with Social-safety-awareness: a reinforcement learning approach from 2D laser scans
- Active Learning in Robotics: A Review of Control Principles
- Reinforcement Learning-Empowered Mobile Edge Computing for 6G Edge Intelligence
- Deep Reinforcement Learning for Radio Resource Allocation and Management in Next Generation Heterogeneous Wireless Networks: A Survey
- DRLinFluids -- An open-source python platform of coupling Deep Reinforcement Learning and OpenFOAM
- ChainerRL: A Deep Reinforcement Learning Library
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Energy-Efficient Thermal Comfort Control in Smart Buildings via Deep Reinforcement Learning
- Hybrid Car-Following Strategy based on Deep Deterministic Policy Gradient and Cooperative Adaptive Cruise Control
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
- Adaptive dynamic programming for nonaffine nonlinear optimal control problem with state constraints
- Hierarchical visuomotor control of humanoids
- Latent Space Policies for Hierarchical Reinforcement Learning
- 3D Simulation for Robot Arm Control with Deep Q-Learning
- Intrinsically motivated reinforcement learning for human-robot interaction in the real-world
- Open-Sourced Reinforcement Learning Environments for Surgical Robotics
- Some Considerations on Learning to Explore via Meta-Reinforcement Learning
- Microswimmers learning chemotaxis with genetic algorithms
- Assessing Transferability from Simulation to Reality for Reinforcement Learning
- Twin actor twin delayed deep deterministic policy gradient (TATD3) learning for batch process control
- Action Robust Reinforcement Learning and Applications in Continuous Control
- A Reinforcement Learning-based Economic Model Predictive Control Framework for Autonomous Operation of Chemical Reactors
- Combining Deep Reinforcement Learning and Safety Based Control for Autonomous Driving
- Discrete Sequential Prediction of Continuous Actions for Deep RL
- Robust Imitation of Diverse Behaviors
- Model-based Deep Reinforcement Learning for Dynamic Portfolio Optimization
- Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning
- Discovering Reinforcement Learning Algorithms
- Sharing Knowledge in Multi-Task Deep Reinforcement Learning
- Multi-agent Reinforcement Learning for Networked System Control
- How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments
- Towards Characterizing Divergence in Deep Q-Learning
- Deep Hierarchical Reinforcement Learning Algorithm in Partially Observable Markov Decision Processes
- Development of a Soft Actor Critic Deep Reinforcement Learning Approach for Harnessing Energy Flexibility in a Large Office Building
- Multi-Task Reinforcement Learning with Soft Modularization
- Skew-Fit: State-Covering Self-Supervised Reinforcement Learning
- Near-Optimal Representation Learning for Hierarchical Reinforcement Learning
- A Survey of Deep Learning for Scientific Discovery
- Learning to Schedule Communication in Multi-agent Reinforcement Learning
- A Self-adaptive SAC-PID Control Approach based on Reinforcement Learning for Mobile Robots
- Deep Reinforcement Learning for Robotic Manipulation-The state of the art
- A Perspective on Deep Learning for Molecular Modeling and Simulations
- Shapley Counterfactual Credits for Multi-Agent Reinforcement Learning
- Deep Imitative Models for Flexible Inference, Planning, and Control
- Language as a Cognitive Tool to Imagine Goals in Curiosity-Driven Exploration
- An experimental evaluation of Deep Reinforcement Learning algorithms for HVAC control
- A Reinforcement Learning approach for Quantum State Engineering
- Testing and verification of neural-network-based safety-critical control software: A systematic literature review
- Multiagent Soft Q-Learning
- Deep reinforcement learning for large-eddy simulation modeling in wall-bounded turbulence
- Data-driven control of room temperature and bidirectional EV charging using deep reinforcement learning: simulations and experiments
- AlphaX: eXploring Neural Architectures with Deep Neural Networks and Monte Carlo Tree Search
- Accelerating Reinforcement Learning for Reaching using Continuous Curriculum Learning
- Modelling the Dynamic Joint Policy of Teammates with Attention Multi-agent DDPG
- GeneraLight: Improving Environment Generalization of Traffic Signal Control via Meta Reinforcement Learning
- Parameter Sharing Deep Deterministic Policy Gradient for Cooperative Multi-agent Reinforcement Learning
- Indirect and Direct Training of Spiking Neural Networks for End-to-End Control of a Lane-Keeping Vehicle
- Training Neural Networks Using Features Replay
- Characterizing Attacks on Deep Reinforcement Learning
- Image quality assessment for machine learning tasks using meta-reinforcement learning
- Deep Imitation Learning for Complex Manipulation Tasks from Virtual Reality Teleoperation
- Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?
- Reinforcement Learning for Generative AI: State of the Art, Opportunities and Open Research Challenges
- Experience-driven Networking: A Deep Reinforcement Learning based Approach
- Learning Scheduling Algorithms for Data Processing Clusters
- Learning End-to-end Multimodal Sensor Policies for Autonomous Navigation
- Data-driven control of micro-climate in buildings: an event-triggered reinforcement learning approach
- An empirical investigation of the challenges of real-world reinforcement learning
- Phasic Policy Gradient
- GenDICE: Generalized Offline Estimation of Stationary Values
- HAQ: Hardware-Aware Automated Quantization with Mixed Precision
- A reinforcement learning approach to rare trajectory sampling
- Improving Coordination in Small-Scale Multi-Agent Deep Reinforcement Learning through Memory-driven Communication
- Reinforcement Learning for Digital Quantum Simulation
- Modified DDPG car-following model with a real-world human driving experience with CARLA simulator
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- A Benchmark Environment Motivated by Industrial Control Problems
- Softmax Deep Double Deterministic Policy Gradients
- Robust Reinforcement Learning on State Observations with Learned Optimal Adversary
- Reinforcement Learning Applications
- AutoGAN-based Dimension Reduction for Privacy Preservation
- Longitudinal Dynamic versus Kinematic Models for Car-Following Control Using Deep Reinforcement Learning
- Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation
- Deep Reinforcement Learning for Autonomous Driving
- Modern Machine Learning Tools for Monitoring and Control of Industrial Processes: A Survey
- Off-Policy Policy Gradient with State Distribution Correction
- A Survey of Optimization Methods from a Machine Learning Perspective
- Model-Free Control for Distributed Stream Data Processing using Deep Reinforcement Learning
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- iGibson 1.0: a Simulation Environment for Interactive Tasks in Large Realistic Scenes
- End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks
- FlingBot: The Unreasonable Effectiveness of Dynamic Manipulation for Cloth Unfolding
- Preparing for the Unknown: Learning a Universal Policy with Online System Identification
- d3rlpy: An Offline Deep Reinforcement Learning Library
- Designing Ecosystems of Intelligence from First Principles
- OnSlicing: Online End-to-End Network Slicing with Reinforcement Learning
- Reinforcement Learning with Prototypical Representations
- CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
- Measuring the Reliability of Reinforcement Learning Algorithms
- Relative Entropy Regularized Policy Iteration
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Deep Reinforcement Learning of Cell Movement in the Early Stage of C. elegans Embryogenesis
- Diversity Policy Gradient for Sample Efficient Quality-Diversity Optimization
- Curiosity-Driven Experience Prioritization via Density Estimation
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- CM3: Cooperative Multi-goal Multi-stage Multi-agent Reinforcement Learning
- Learning to Paint With Model-based Deep Reinforcement Learning
- Model-Augmented Actor-Critic: Backpropagating through Paths
- Hierarchical Reinforcement Learning with Hindsight
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
- Constrained Attractor Selection Using Deep Reinforcement Learning
- Model-Free Mean-Field Reinforcement Learning: Mean-Field MDP and Mean-Field Q-Learning
- ACCNet: Actor-Coordinator-Critic Net for "Learning-to-Communicate" with Deep Multi-agent Reinforcement Learning
- Path Integral Networks: End-to-End Differentiable Optimal Control
- Bayesian Sequential Optimal Experimental Design for Nonlinear Models Using Policy Gradient Reinforcement Learning
- Dissipative SymODEN: Encoding Hamiltonian Dynamics with Dissipation and Control into Deep Learning
- Search on the Replay Buffer: Bridging Planning and Reinforcement Learning
- Safety and Liveness Guarantees through Reach-Avoid Reinforcement Learning
- Primal Wasserstein Imitation Learning
- Reinforcement Learning to Rank in E-Commerce Search Engine: Formalization, Analysis, and Application
- Reinforcement Learning and Deep Learning based Lateral Control for Autonomous Driving
- Learning Permutations with Sinkhorn Policy Gradient
- Towards a Common Implementation of Reinforcement Learning for Multiple Robotic Tasks
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State Observations
- DeepTraffic: Crowdsourced Hyperparameter Tuning of Deep Reinforcement Learning Systems for Multi-Agent Dense Traffic Navigation
- A Comprehensive Overview and Survey of Recent Advances in Meta-Learning
- Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation
- Transforming Cooling Optimization for Green Data Center via Deep Reinforcement Learning
- Inclined Quadrotor Landing using Deep Reinforcement Learning
- Global Convergence of Policy Gradient Methods to (Almost) Locally Optimal Policies
- Note on Attacking Object Detectors with Adversarial Stickers
- Hierarchical Deep Multiagent Reinforcement Learning with Temporal Abstraction
- Towards Understanding Asynchronous Advantage Actor-critic: Convergence and Linear Speedup
- Reinforcement Learning for Robotic Manipulation using Simulated Locomotion Demonstrations
- Deep Reinforcement Learning with Population-Coded Spiking Neural Network for Continuous Control
- Grounded Language Learning Fast and Slow
- Hindsight policy gradients
- Lipschitz Continuity in Model-based Reinforcement Learning
- Deep Multi-agent Reinforcement Learning for Highway On-Ramp Merging in Mixed Traffic
- Reinforcement Learning for Pivoting Task
- VRKitchen: an Interactive 3D Virtual Environment for Task-oriented Learning
- Hyp-RL : Hyperparameter Optimization by Reinforcement Learning
- Deep Multi-Agent Reinforcement Learning with Relevance Graphs
- Automatic Curriculum Learning through Value Disagreement
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
- TD-Regularized Actor-Critic Methods
- Recurrent Deterministic Policy Gradient Method for Bipedal Locomotion on Rough Terrain Challenge
- Deep Reinforcement Learning with Successor Features for Navigation across Similar Environments
- Towards Efficient Training for Neural Network Quantization
- Learning Deployable Navigation Policies at Kilometer Scale from a Single Traversal
- Active Domain Randomization
- Prioritized Sequence Experience Replay
- Deep Reinforcement Learning with a Natural Language Action Space
- Overcoming Exploration: Deep Reinforcement Learning for Continuous Control in Cluttered Environments from Temporal Logic Specifications
- Affect-Driven Modelling of Robot Personality for Collaborative Human-Robot Interactions
- Pretraining Deep Actor-Critic Reinforcement Learning Algorithms With Expert Demonstrations
- If MaxEnt RL is the Answer, What is the Question?
- Deep Reinforcement Learning for Six Degree-of-Freedom Planetary Powered Descent and Landing
- Coordinated Exploration via Intrinsic Rewards for Multi-Agent Reinforcement Learning
- RLCard: A Toolkit for Reinforcement Learning in Card Games
- Computational Theories of Curiosity-Driven Learning
- Crafting a Toolchain for Image Restoration by Deep Reinforcement Learning
- Goal-conditioned Imitation Learning
- Logarithmic Regret for Adversarial Online Control
- A Reinforcement Learning Approach for Transient Control of Liquid Rocket Engines
- Deep Reinforcement Learning from Policy-Dependent Human Feedback
- Naive Exploration is Optimal for Online LQR
- BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning
- Cooperative Behavior Planning for Automated Driving using Graph Neural Networks
- MSPM: A Modularized and Scalable Multi-Agent Reinforcement Learning-based System for Financial Portfolio Management
- The Evolution of Reinforcement Learning in Quantitative Finance: A Survey
- Autonomous Drone Swarm Navigation and Multi-target Tracking in 3D Environments with Dynamic Obstacles
- Interactive Differentiable Simulation
- Model-free Deep Reinforcement Learning for Urban Autonomous Driving
- Natural Environment Benchmarks for Reinforcement Learning
- Mapping State Space using Landmarks for Universal Goal Reaching
- Trajectory Planning with Deep Reinforcement Learning in High-Level Action Spaces
- Sim-to-Real Transfer of Accurate Grasping with Eye-In-Hand Observations and Continuous Control
- A Tutorial on Ultra-Reliable and Low-Latency Communications in 6G: Integrating Domain Knowledge into Deep Learning
- Diff-DAC: Distributed Actor-Critic for Average Multitask Deep Reinforcement Learning
- Classification with Costly Features as a Sequential Decision-Making Problem
- Automated Cloud Provisioning on AWS using Deep Reinforcement Learning
- VIREL: A Variational Inference Framework for Reinforcement Learning
- WD3: Taming the Estimation Bias in Deep Reinforcement Learning
- Deep Reinforcement Learning for High Precision Assembly Tasks
- Parameterized Reinforcement Learning for Optical System Optimization
- Dynamic Adversarial Patch for Evading Object Detection Models
- Constrained Upper Confidence Reinforcement Learning
- Learning Shape Control of Elastoplastic Deformable Linear Objects
- Learning Robotic Navigation from Experience: Principles, Methods, and Recent Results
- Maximum Entropy RL (Provably) Solves Some Robust RL Problems
- Multi-Agent Reinforcement Learning for Dynamic Ocean Monitoring by a Swarm of Buoys
- Generalizable control for multiparameter quantum metrology
- MADRaS : Multi Agent Driving Simulator
- Continuous-Discrete Reinforcement Learning for Hybrid Control in Robotics
- Model-Based Safe Reinforcement Learning with Time-Varying State and Control Constraints: An Application to Intelligent Vehicles
- Trust-PCL: An Off-Policy Trust Region Method for Continuous Control
- A Survey of Deep Learning Techniques for Mobile Robot Applications
- Autonomous Driving in Reality with Reinforcement Learning and Image Translation
- Constructing Parsimonious Analytic Models for Dynamic Systems via Symbolic Regression
- Deep reinforcement learning in World-Earth system models to discover sustainable management strategies
- Deep Reinforcement Learning Methods for Structure-Guided Processing Path Optimization
- Learning to Design Circuits
- Trust-Region Method with Deep Reinforcement Learning in Analog Design Space Exploration
- Randomized Ensembled Double Q-Learning: Learning Fast Without a Model
- Energy-Based Hindsight Experience Prioritization
- The Ingredients of Real-World Robotic Reinforcement Learning
- Learning to Reach Goals via Iterated Supervised Learning
- Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination
- Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network
- Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
- Neuroflight: Next Generation Flight Control Firmware
- IPAPRec: A promising tool for learning high-performance mapless navigation skills with deep reinforcement learning
- DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
- An Adaptive Clipping Approach for Proximal Policy Optimization
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- When Cyber-Physical Systems Meet AI: A Benchmark, an Evaluation, and a Way Forward
- Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning
- Learning Functionally Decomposed Hierarchies for Continuous Control Tasks with Path Planning
- Joint Multi-Dimension Pruning via Numerical Gradient Update
- Hierarchical Reinforcement Learning By Discovering Intrinsic Options
- Safe reinforcement learning for probabilistic reachability and safety specifications: A Lyapunov-based approach
- A unified strategy for implementing curiosity and empowerment driven reinforcement learning
- Learning to Explore with Meta-Policy Gradient
- Off-Policy Evaluation via Off-Policy Classification
- Multi-task Deep Reinforcement Learning with PopArt
- Pre-training with Non-expert Human Demonstration for Deep Reinforcement Learning
- TriFinger: An Open-Source Robot for Learning Dexterity
- Accuracy-based Curriculum Learning in Deep Reinforcement Learning
- Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control
- Interpretable End-to-end Urban Autonomous Driving with Latent Deep Reinforcement Learning
- Learning Event-triggered Control from Data through Joint Optimization
- Learning to Fly via Deep Model-Based Reinforcement Learning
- A Reinforcement Learning Approach for the Continuous Electricity Market of Germany: Trading from the Perspective of a Wind Park Operator
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEs
- Benchmarks for Deep Off-Policy Evaluation
- Cooperative Multi-Agent Reinforcement Learning for Low-Level Wireless Communication
- A Deep Reinforcement Learning Framework for Rebalancing Dockless Bike Sharing Systems
- Behavioral decision-making for urban autonomous driving in the presence of pedestrians using Deep Recurrent Q-Network
- Learning Self-Imitating Diverse Policies
- Behavioural Repertoire via Generative Adversarial Policy Networks
- Learning 6-DoF Grasping and Pick-Place Using Attention Focus
- Forward and inverse reinforcement learning sharing network weights and hyperparameters
- Deep hierarchical reinforcement agents for automated penetration testing
- Combining Neural Networks and Tree Search for Task and Motion Planning in Challenging Environments
- Mirror Descent Policy Optimization
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- World Model as a Graph: Learning Latent Landmarks for Planning
- Continual Reinforcement Learning with Complex Synapses
- Video Captioning via Hierarchical Reinforcement Learning
- Towards Interpretable-AI Policies Induction using Evolutionary Nonlinear Decision Trees for Discrete Action Systems
- Deep Reinforcement Learning for Unmanned Aerial Vehicle-Assisted Vehicular Networks
- Reducing Bus Bunching with Asynchronous Multi-Agent Reinforcement Learning
- Learning Latent Plans from Play
- Modular Deep Reinforcement Learning with Temporal Logic Specifications
- Sample-efficient Reinforcement Learning Representation Learning with Curiosity Contrastive Forward Dynamics Model
- Optimal Stroke Learning with Policy Gradient Approach for Robotic Table Tennis
- Offline Reinforcement Learning with Reverse Model-based Imagination
- A Tour of Reinforcement Learning: The View from Continuous Control
- Sample-Efficient Imitation Learning via Generative Adversarial Nets
- Energy-Efficient Control Adaptation with Safety Guarantees for Learning-Enabled Cyber-Physical Systems
- Cooperative Lane Changing via Deep Reinforcement Learning
- Learning to Assist Drone Landings
- Quality of service based radar resource management using deep reinforcement learning
- Feudal Multi-Agent Hierarchies for Cooperative Reinforcement Learning
- On-Policy Robot Imitation Learning from a Converging Supervisor
- A Deep Recurrent Q Network towards Self-adapting Distributed Microservices architecture
- A Centralised Soft Actor Critic Deep Reinforcement Learning Approach to District Demand Side Management through CityLearn
- TF-Replicator: Distributed Machine Learning for Researchers
- Fresh, Fair and Energy-Efficient Content Provision in a Private and Cache-Enabled UAV Network
- Deep Reinforcement Learning for Human-Like Driving Policies in Collision Avoidance Tasks of Self-Driving Cars
- Stability-certified reinforcement learning: A control-theoretic perspective
- Deep Intrinsically Motivated Continuous Actor-Critic for Efficient Robotic Visuomotor Skill Learning
- Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
- State Entropy Maximization with Random Encoders for Efficient Exploration
- Generalizable Episodic Memory for Deep Reinforcement Learning
- Accelerating Quadratic Optimization with Reinforcement Learning
- Bipedal Walking Robot using Deep Deterministic Policy Gradient
- Exploration versus exploitation in reinforcement learning: a stochastic control approach
- Reinforcement Learning of Active Vision for Manipulating Objects under Occlusions
- Reinforced Neighborhood Selection Guided Multi-Relational Graph Neural Networks
- The Faults in Our Pi Stars: Security Issues and Open Challenges in Deep Reinforcement Learning
- Leveraging the Capabilities of Connected and Autonomous Vehicles and Multi-Agent Reinforcement Learning to Mitigate Highway Bottleneck Congestion
- Recovery RL: Safe Reinforcement Learning with Learned Recovery Zones
- Formulation and validation of a car-following model based on deep reinforcement learning
- VisuoSpatial Foresight for Multi-Step, Multi-Task Fabric Manipulation
- Reinforcement Learning on Variable Impedance Controller for High-Precision Robotic Assembly
- Solving NP-Hard Problems on Graphs with Extended AlphaGo Zero
- A Survey of Deep Reinforcement Learning in Recommender Systems: A Systematic Review and Future Directions
- Distributed multi-agent target search and tracking with Gaussian process and reinforcement learning
- A Practical Guide to Multi-Objective Reinforcement Learning and Planning
- Model-Based Offline Planning
- Online Meta-Critic Learning for Off-Policy Actor-Critic Methods
- Tsallis Reinforcement Learning: A Unified Framework for Maximum Entropy Reinforcement Learning
- Reinforcement Learning for Photonic Component Design
- Optimizing Throughput Performance in Distributed MIMO Wi-Fi Networks using Deep Reinforcement Learning
- Deep Reinforcement Learning for Dynamic Multichannel Access in Wireless Networks
- Dropout Q-Functions for Doubly Efficient Reinforcement Learning
- Learning a Decentralized Multi-arm Motion Planner
- An Open-Source Multi-Goal Reinforcement Learning Environment for Robotic Manipulation with Pybullet
- PlanGAN: Model-based Planning With Sparse Rewards and Multiple Goals
- SafeCritic: Collision-Aware Trajectory Prediction
- Value Functions Factorization with Latent State Information Sharing in Decentralized Multi-Agent Policy Gradients
- The problem with DDPG: understanding failures in deterministic environments with sparse rewards
- Differential Variable Speed Limits Control for Freeway Recurrent Bottlenecks via Deep Reinforcement learning
- Meta-SAC: Auto-tune the Entropy Temperature of Soft Actor-Critic via Metagradient
- A Survey on Autonomous Vehicle Control in the Era of Mixed-Autonomy: From Physics-Based to AI-Guided Driving Policy Learning
- Quantum reinforcement learning in continuous action space
- Quantile QT-Opt for Risk-Aware Vision-Based Robotic Grasping
- Imitation Learning for High Precision Peg-in-Hole Tasks
- Distributionally Robust Reinforcement Learning
- Optimal Attacks on Reinforcement Learning Policies
- Experience Replay with Likelihood-free Importance Weights
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
- Training Agents using Upside-Down Reinforcement Learning
- Feedback Linearization for Unknown Systems via Reinforcement Learning
- Sparse Graphical Memory for Robust Planning
- Online Data Poisoning Attack
- Deep reinforcement learning for smart calibration of radio telescopes
- Iroko: A Framework to Prototype Reinforcement Learning for Data Center Traffic Control
- Learning to Manipulate Deformable Objects without Demonstrations
- A View on Deep Reinforcement Learning in System Optimization
- Human-like Energy Management Based on Deep Reinforcement Learning and Historical Driving Experiences
- Learning Robotic Manipulation through Visual Planning and Acting
- Generalized Hindsight for Reinforcement Learning
- AC-Teach: A Bayesian Actor-Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- Applications of Deep Reinforcement Learning in Communications and Networking: A Survey
- Reinforcement Learning for Shared Autonomy Drone Landings
- A physics-informed reinforcement learning approach for the interfacial area transport in two-phase flow
- Reinforcement Learning for Multi-Product Multi-Node Inventory Management in Supply Chains
- Behavior Self-Organization Supports Task Inference for Continual Robot Learning
- AutoPrivacy: Automated Layer-wise Parameter Selection for Secure Neural Network Inference
- Delay-Aware Multi-Agent Reinforcement Learning for Cooperative and Competitive Environments
- Regularization Matters in Policy Optimization
- Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning
- Dynamics-aware Embeddings
- QTRAN++: Improved Value Transformation for Cooperative Multi-Agent Reinforcement Learning
- Self-Supervised Exploration via Disagreement
- Learning to Run with Actor-Critic Ensemble
- Seeing isn't Believing: Practical Adversarial Attack Against Object Detectors
- Regularized Softmax Deep Multi-Agent -Learning
- VR-Goggles for Robots: Real-to-sim Domain Adaptation for Visual Control
- Learning to Combat Compounding-Error in Model-Based Reinforcement Learning
- Partially Observable Markov Decision Process for Recommender Systems
- Deep Reinforcement Learning Based Dynamic Trajectory Control for UAV-assisted Mobile Edge Computing
- Dynamic Programming Principles for Mean-Field Controls with Learning
- Learning to Search via Retrospective Imitation
- DRLDO: A novel DRL based De-ObfuscationSystem for Defense against Metamorphic Malware
- Deep Reinforcement Learning with Robust and Smooth Policy
- Improving Robot Dual-System Motor Learning with Intrinsically Motivated Meta-Control and Latent-Space Experience Imagination
- Accelerating Reinforcement Learning through GPU Atari Emulation
- Learning to Factor Policies and Action-Value Functions: Factored Action Space Representations for Deep Reinforcement learning
- A Unified Bellman Optimality Principle Combining Reward Maximization and Empowerment
- Towards Cooperation in Sequential Prisoner's Dilemmas: a Deep Multiagent Reinforcement Learning Approach
- Safety-Guided Deep Reinforcement Learning via Online Gaussian Process Estimation
- Hardware-Centric AutoML for Mixed-Precision Quantization
- On the model-based stochastic value gradient for continuous reinforcement learning
- Deep Predictive Policy Training using Reinforcement Learning
- Multi-Agent Constrained Policy Optimisation
- Local policy search with Bayesian optimization
- Deep Reinforcement Learning for Contact-Rich Skills Using Compliant Movement Primitives
- Hierarchical Reinforcement Learning Framework for Stochastic Spaceflight Campaign Design
- Flight Controller Synthesis Via Deep Reinforcement Learning
- Hamilton-Jacobi Deep Q-Learning for Deterministic Continuous-Time Systems with Lipschitz Continuous Controls
- From self-tuning regulators to reinforcement learning and back again
- Boredom-driven curious learning by Homeo-Heterostatic Value Gradients
- Goal-directed graph construction using reinforcement learning
- Deep Reinforcement Learning for Wireless Scheduling in Distributed Networked Control
- Estimation Error Correction in Deep Reinforcement Learning for Deterministic Actor-Critic Methods
- Robust Deep Reinforcement Learning through Adversarial Loss
- Deep Reinforcement Learning for Dexterous Manipulation with Concept Networks
- Generation and storage of spin squeezing via learning-assisted optimal control
- Guided Policy Search as Approximate Mirror Descent
- Deep reinforcement learning for robust quantum optimization
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
- Can Increasing Input Dimensionality Improve Deep Reinforcement Learning?
- Learning Nash Equilibrium for General-Sum Markov Games from Batch Data
- Deep Reinforcement Learning for Task Offloading in Mobile Edge Computing Systems
- Reinforcement Learning with Perturbed Rewards
- Deep Reinforcement Learning: Framework, Applications, and Embedded Implementations
- Deep Reinforcement Learning for Event-Triggered Control
- RTFM: Generalising to Novel Environment Dynamics via Reading
- Motion Planner Augmented Reinforcement Learning for Robot Manipulation in Obstructed Environments
- Hindsight Goal Ranking on Replay Buffer for Sparse Reward Environment
- Modern Deep Reinforcement Learning Algorithms
- Precision medicine as a control problem: Using simulation and deep reinforcement learning to discover adaptive, personalized multi-cytokine therapy for sepsis
- Online Off-policy Prediction
- Goal-Auxiliary Actor-Critic for 6D Robotic Grasping with Point Clouds
- Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks
- Remember and Forget for Experience Replay
- Stochastic Recursive Momentum for Policy Gradient Methods
- Learning Visual Robotic Control Efficiently with Contrastive Pre-training and Data Augmentation
- A Hierarchical Architecture for Sequential Decision-Making in Autonomous Driving using Deep Reinforcement Learning
- Bayesian Policy Gradients via Alpha Divergence Dropout Inference
- Randomized Value Functions via Multiplicative Normalizing Flows
- Policy Evaluation Networks
- Toward Packet Routing with Fully-distributed Multi-agent Deep Reinforcement Learning
- Promoting Coordination through Policy Regularization in Multi-Agent Deep Reinforcement Learning
- Discretizing Continuous Action Space for On-Policy Optimization
- Multi-agent Reinforcement Learning Accelerated MCMC on Multiscale Inversion Problem
- Power Allocation in Multi-User Cellular Networks: Deep Reinforcement Learning Approaches
- Learning to Collaborate: Multi-Scenario Ranking via Multi-Agent Reinforcement Learning
- Learning Multi-Agent Coordination through Connectivity-driven Communication
- Model-based Reinforcement Learning from Signal Temporal Logic Specifications
- Trainable Greedy Decoding for Neural Machine Translation
- Transfer learning from synthetic to real images using variational autoencoders for robotic applications
- Recurrent Off-policy Baselines for Memory-based Continuous Control
- Applications of Multi-Agent Reinforcement Learning in Future Internet: A Comprehensive Survey
- Internal Model from Observations for Reward Shaping
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling
- Deep Reinforcement Learning for Motion Planning of Mobile Robots
- Using Reinforcement Learning to find Efficient Qubit Routing Policies for Deployment in Near-term Quantum Computers
- On the Estimation Bias in Double Q-Learning
- Online 3D Bin Packing with Constrained Deep Reinforcement Learning
- Deep Reinforcement Learning for Autonomous Internet of Things: Model, Applications and Challenges
- Tactical Optimism and Pessimism for Deep Reinforcement Learning
- MHER: Model-based Hindsight Experience Replay
- StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction
- Statistical Bootstrapping for Uncertainty Estimation in Off-Policy Evaluation
- Observation Space Matters: Benchmark and Optimization Algorithm
- SCARL: Side-Channel Analysis with Reinforcement Learning on the Ascon Authenticated Cipher
- DAQN: Deep Auto-encoder and Q-Network
- Deep reinforcement learning under signal temporal logic constraints using Lagrangian relaxation
- Ctrl-Z: Recovering from Instability in Reinforcement Learning
- Hyperparameter Optimization via Sequential Uniform Designs
- Model-based Reinforcement Learning for Predictions and Control for Limit Order Books
- An initial attempt of combining visual selective attention with deep reinforcement learning
- Reinforcement co-Learning of Deep and Spiking Neural Networks for Energy-Efficient Mapless Navigation with Neuromorphic Hardware
- Learning Actionable Representations from Visual Observations
- PLAS: Latent Action Space for Offline Reinforcement Learning
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Constraint Learning for Control Tasks with Limited Duration Barrier Functions
- Curiosity-driven reinforcement learning with homeostatic regulation
- Adversarial Evaluation of Autonomous Vehicles in Lane-Change Scenarios
- Neural Temporal-Difference and Q-Learning Provably Converge to Global Optima
- Action Semantics Network: Considering the Effects of Actions in Multiagent Systems
- Gated networks: an inventory
- Path Integral Guided Policy Search
- Fixed-Dimensional and Permutation Invariant State Representation of Autonomous Driving
- F2A2: Flexible Fully-decentralized Approximate Actor-critic for Cooperative Multi-agent Reinforcement Learning
- Real-Time Model Calibration with Deep Reinforcement Learning
- Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video
- Intervention Aided Reinforcement Learning for Safe and Practical Policy Optimization in Navigation
- A Boolean Task Algebra for Reinforcement Learning
- Attractor Selection in Nonlinear Energy Harvesting Using Deep Reinforcement Learning
- Learning Heuristic Search via Imitation
- A Survey of Reinforcement Learning Techniques: Strategies, Recent Development, and Future Directions
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance Guarantees
- A Meta-Reinforcement Learning Approach to Process Control
- Depth Control of Model-Free AUVs via Reinforcement Learning
- Robust Reinforcement Learning via Adversarial training with Langevin Dynamics
- Entropy Regularization for Mean Field Games with Learning
- Upper Confidence Primal-Dual Reinforcement Learning for CMDP with Adversarial Loss
- Learning Fast Adaptation with Meta Strategy Optimization
- Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations
- How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy Optimization
- Importance mixing: Improving sample reuse in evolutionary policy search methods
- Self-Imitation Learning via Generalized Lower Bound Q-learning
- Neural Replicator Dynamics
- Deep reinforcement learning approach to MIMO precoding problem: Optimality and Robustness
- Lifelong Policy Gradient Learning of Factored Policies for Faster Training Without Forgetting
- Counterfactual Data Augmentation using Locally Factored Dynamics
- Rapidly Adaptable Legged Robots via Evolutionary Meta-Learning
- Using Simulation Optimization to Improve Zero-shot Policy Transfer of Quadrotors
- Reinforced Epidemic Control: Saving Both Lives and Economy
- Microscopic Traffic Simulation by Cooperative Multi-agent Deep Reinforcement Learning
- Curiosity-Driven Multi-Criteria Hindsight Experience Replay
- Automata Guided Reinforcement Learning With Demonstrations
- Provably Robust Blackbox Optimization for Reinforcement Learning
- Adversarial Imitation Learning via Random Search
- Learning Multi-agent Communication under Limited-bandwidth Restriction for Internet Packet Routing
- Contrastive Explanations for Reinforcement Learning via Embedded Self Predictions
- Optimistic Policy Optimization with Bandit Feedback
- Ablation of a Robot's Brain: Neural Networks Under a Knife
- Challenges of Applying Deep Reinforcement Learning in Dynamic Dispatching
- Automatic Inverse Treatment Planning for Gamma Knife Radiosurgery via Deep Reinforcement Learning
- Automated vehicle's behavior decision making using deep reinforcement learning and high-fidelity simulation environment
- Dynamic Pricing on E-commerce Platform with Deep Reinforcement Learning: A Field Experiment
- An Open-Source Framework for Adaptive Traffic Signal Control
- ROBEL: Robotics Benchmarks for Learning with Low-Cost Robots
- Case Study: Verifying the Safety of an Autonomous Racing Car with a Neural Network Controller
- Leveraging exploration in off-policy algorithms via normalizing flows
- Self-Adversarial Learning with Comparative Discrimination for Text Generation
- Safe-visor Architecture for Sandboxing (AI-based) Unverified Controllers in Stochastic Cyber-Physical Systems
- Soft Hindsight Experience Replay
- MADE: Exploration via Maximizing Deviation from Explored Regions
- Imitation Learning: Progress, Taxonomies and Challenges
- Positive-Unlabeled Reward Learning
- Safe Deep Reinforcement Learning for Multi-Agent Systems with Continuous Action Spaces
- Self-supervised Learning of Distance Functions for Goal-Conditioned Reinforcement Learning
- Revisiting Parameter Sharing in Multi-Agent Deep Reinforcement Learning
- Learning in Nonzero-Sum Stochastic Games with Potentials
- Learning and Planning in Complex Action Spaces
- Smoothed Action Value Functions for Learning Gaussian Policies
- Image quality assessment by overlapping task-specific and task-agnostic measures: application to prostate multiparametric MR images for cancer segmentation
- Large-scale Interactive Recommendation with Tree-structured Policy Gradient
- Composing Task-Agnostic Policies with Deep Reinforcement Learning
- Map-based Multi-Policy Reinforcement Learning: Enhancing Adaptability of Robots by Deep Reinforcement Learning
- Neural Architecture Search using Deep Neural Networks and Monte Carlo Tree Search
- A Re-classification of Information Seeking Tasks and Their Computational Solutions
- Snooping Attacks on Deep Reinforcement Learning
- Off-policy Maximum Entropy Reinforcement Learning : Soft Actor-Critic with Advantage Weighted Mixture Policy(SAC-AWMP)
- Learning to Multi-Task by Active Sampling
- Hamilton-Jacobi-Bellman Equations for Q-Learning in Continuous Time
- Sample-Efficient Learning of Nonprehensile Manipulation Policies via Physics-Based Informed State Distributions
- Conditional Driving from Natural Language Instructions
- Robotic Surgery With Lean Reinforcement Learning
- Learning Agent Communication under Limited Bandwidth by Message Pruning
- Actor-Critic Reinforcement Learning for Control with Stability Guarantee
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience Replay
- Multi-task Batch Reinforcement Learning with Metric Learning
- Winning Isn't Everything: Enhancing Game Development with Intelligent Agents
- ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning
- State-only Imitation with Transition Dynamics Mismatch
- Deep Reinforcement Learning in a Monetary Model
- Dynamic Experience Replay
- On the Weaknesses of Reinforcement Learning for Neural Machine Translation
- Emergent Real-World Robotic Skills via Unsupervised Off-Policy Reinforcement Learning
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Safe Reinforcement Learning of Control-Affine Systems with Vertex Networks
- Causality and Batch Reinforcement Learning: Complementary Approaches To Planning In Unknown Domains
- Integrating Behavior Cloning and Reinforcement Learning for Improved Performance in Dense and Sparse Reward Environments
- Causal Influence Detection for Improving Efficiency in Reinforcement Learning
- Deep Reinforcement Learning for Dynamic Treatment Regimes on Medical Registry Data
- Directed Exploration for Reinforcement Learning
- Learning Convex Optimization Control Policies
- Enforcing constraints for time series prediction in supervised, unsupervised and reinforcement learning
- Continual Reinforcement Learning with Multi-Timescale Replay
- Model-based Lookahead Reinforcement Learning
- Learning to Collaborate in Multi-Module Recommendation via Multi-Agent Reinforcement Learning without Communication
- Deep Actor-Critic Learning for Distributed Power Control in Wireless Mobile Networks
- Policy Gradient in Partially Observable Environments: Approximation and Convergence
- Optimizing Mixed Autonomy Traffic Flow With Decentralized Autonomous Vehicles and Multi-Agent RL
- Uncertainty-aware Model-based Policy Optimization
- Symbolic Regression Methods for Reinforcement Learning
- Offline Reinforcement Learning for Autonomous Driving with Safety and Exploration Enhancement
- Design Space of Behaviour Planning for Autonomous Driving
- Towards Effective Context for Meta-Reinforcement Learning: an Approach based on Contrastive Learning
- Large-scale traffic signal control using machine learning: some traffic flow considerations
- Optimization Landscape of Gradient Descent for Discrete-time Static Output Feedback
- Semantic-Transferable Weakly-Supervised Endoscopic Lesions Segmentation
- Dependability Analysis of Deep Reinforcement Learning based Robotics and Autonomous Systems through Probabilistic Model Checking
- Neural Stochastic Dual Dynamic Programming
- Comparing Deep Reinforcement Learning and Evolutionary Methods in Continuous Control
- End-to-end driving simulation via angle branched network
- Chance-Constrained Control with Lexicographic Deep Reinforcement Learning
- Attention Control with Metric Learning Alignment for Image Set-based Recognition
- Offline Meta Learning of Exploration
- The Emergence of Wireless MAC Protocols with Multi-Agent Reinforcement Learning
- Hardware as Policy: Mechanical and Computational Co-Optimization using Deep Reinforcement Learning
- ADVERSARIALuscator: An Adversarial-DRL Based Obfuscator and Metamorphic Malware SwarmGenerator
- Multi-Agent Actor-Critic with Generative Cooperative Policy Network
- Meta Reinforcement Learning with Task Embedding and Shared Policy
- Risk Adversarial Learning System for Connected and Autonomous Vehicle Charging
- Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads
- Smooth Exploration for Robotic Reinforcement Learning
- Optimistic Bull or Pessimistic Bear: Adaptive Deep Reinforcement Learning for Stock Portfolio Allocation
- Affordance-based Reinforcement Learning for Urban Driving
- Multi-Agent Actor-Critic with Hierarchical Graph Attention Network
- Actor-critic versus direct policy search: a comparison based on sample complexity
- Continuous Motion Planning with Temporal Logic Specifications using Deep Neural Networks
- Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences
- Maximum Mutation Reinforcement Learning for Scalable Control
- A Multi-Agent Reinforcement Learning Method for Impression Allocation in Online Display Advertising
- Sample Complexity of Reinforcement Learning using Linearly Combined Model Ensembles
- Data-Efficient Policy Evaluation Through Behavior Policy Search
- Stabilization of vertical motion of a vehicle on bumpy terrain using deep reinforcement learning
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
- Guided Policy Search Model-based Reinforcement Learning for Urban Autonomous Driving
- Scalable Voltage Control using Structure-Driven Hierarchical Deep Reinforcement Learning
- Geometric Entropic Exploration
- Context-Aware Safe Reinforcement Learning for Non-Stationary Environments
- Deep Reinforcement Learning for Tropical Air Free-Cooled Data Center Control
- Thief, Beware of What Get You There: Towards Understanding Model Extraction Attack
- ACtuAL: Actor-Critic Under Adversarial Learning
- Deep Reinforcement Learning to Acquire Navigation Skills for Wheel-Legged Robots in Complex Environments
- Evolving Inborn Knowledge For Fast Adaptation in Dynamic POMDP Problems
- A Deep Reinforcement Learning Approach for Traffic Signal Control Optimization
- Automating Vehicles by Deep Reinforcement Learning using Task Separation with Hill Climbing
- Reinforcement Learning Your Way: Agent Characterization through Policy Regularization
- Learning Deep Mean Field Games for Modeling Large Population Behavior
- The Effect of Multi-step Methods on Overestimation in Deep Reinforcement Learning
- Effective Exploration for Deep Reinforcement Learning via Bootstrapped Q-Ensembles under Tsallis Entropy Regularization
- Adversarial Reinforcement Learning for Observer Design in Autonomous Systems under Cyber Attacks
- Multi-Agent Deep Reinforcement Learning Based Trajectory Planning for Multi-UAV Assisted Mobile Edge Computing
- Medical Dead-ends and Learning to Identify High-risk States and Treatments
- Trajectory Design for UAV-Based Internet-of-Things Data Collection: A Deep Reinforcement Learning Approach
- Offline Meta-Reinforcement Learning with Online Self-Supervision
- Decentralized Q-Learning in Zero-sum Markov Games
- Value Iteration in Continuous Actions, States and Time
- Capability Iteration Network for Robot Path Planning
- Computational Performance of Deep Reinforcement Learning to find Nash Equilibria
- Efficient time stepping for numerical integration using reinforcement learning
- Global Convergence of Policy Gradient Primal-dual Methods for Risk-constrained LQRs
- Storchastic: A Framework for General Stochastic Automatic Differentiation
- An RL-Based Adaptive Detection Strategy to Secure Cyber-Physical Systems
- Doubly Robust Off-Policy Actor-Critic: Convergence and Optimality
- Derivative-Free Policy Optimization for Linear Risk-Sensitive and Robust Control Design: Implicit Regularization and Sample Complexity
- Deep Reinforcement Learning for Long-Short Portfolio Optimization
- Learning Light Transport the Reinforced Way
- Injecting Prior Knowledge for Transfer Learning into Reinforcement Learning Algorithms using Logic Tensor Networks
- Multimodal Safety-Critical Scenarios Generation for Decision-Making Algorithms Evaluation
- Efficient Connected and Automated Driving System with Multi-agent Graph Reinforcement Learning
- Designing high-fidelity multi-qubit gates for semiconductor quantum dots through deep reinforcement learning
- Learning Hierarchical Teaching Policies for Cooperative Agents
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and Planning
- Delay-Aware Model-Based Reinforcement Learning for Continuous Control
- Imitation-Projected Programmatic Reinforcement Learning
- Hindsight Trust Region Policy Optimization
- Dependency-aware Attention Control for Unconstrained Face Recognition with Image Sets
- Improving Generalization of Reinforcement Learning with Minimax Distributional Soft Actor-Critic
- Long-Term Visitation Value for Deep Exploration in Sparse Reward Reinforcement Learning
- Direct and indirect reinforcement learning
- OffWorld Gym: open-access physical robotics environment for real-world reinforcement learning benchmark and research
- Safety Augmented Value Estimation from Demonstrations (SAVED): Safe Deep Model-Based RL for Sparse Cost Robotic Tasks
- Reinforcement Learning for Portfolio Management
- A Hybrid Stochastic Policy Gradient Algorithm for Reinforcement Learning
- Policy-Aware Model Learning for Policy Gradient Methods
- Resilient Load Restoration in Microgrids Considering Mobile Energy Storage Fleets: A Deep Reinforcement Learning Approach
- Deep Residual Reinforcement Learning
- Hierarchical Reinforcement Learning Method for Autonomous Vehicle Behavior Planning
- EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RL
- Deep Reinforcement Learning for Autonomous Ground Vehicle Exploration Without A-Priori Maps
- Tonic: A Deep Reinforcement Learning Library for Fast Prototyping and Benchmarking
- SmartChoices: Hybridizing Programming and Machine Learning
- Transferring human emotions to robot motions using Neural Policy Style Transfer
- Safe Reinforcement Learning for Autonomous Vehicles through Parallel Constrained Policy Optimization
- Deep Reinforcement Learning for Demand Driven Services in Logistics and Transportation Systems: A Survey
- Optimal Control of Complex Systems through Variational Inference with a Discrete Event Decision Process
- Principled Exploration via Optimistic Bootstrapping and Backward Induction
- Exploratory State Representation Learning
- Policy Optimization With Penalized Point Probability Distance: An Alternative To Proximal Policy Optimization
- Dynamic Control of a Fiber Manufacturing Process using Deep Reinforcement Learning
- Finding the best design parameters for optical nanostructures using reinforcement learning
- Parallel Knowledge Transfer in Multi-Agent Reinforcement Learning
- Vehicular Cooperative Perception Through Action Branching and Federated Reinforcement Learning
- Learning Actionable Representations with Goal-Conditioned Policies
- VisuoSpatial Foresight for Physical Sequential Fabric Manipulation
- A Framework for Studying Reinforcement Learning and Sim-to-Real in Robot Soccer
- SORNet: Spatial Object-Centric Representations for Sequential Manipulation
- Deep Reinforcement Learning Based Optimization for IRS Based UAV-NOMA Downlink Networks
- Local Search for Policy Iteration in Continuous Control
- Can Interpretable Reinforcement Learning Manage Prosperity Your Way?
- Entanglement engineering of optomechanical systems by reinforcement learning
- Variational Adaptive-Newton Method for Explorative Learning
- Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a Survey
- Synthesizing Action Sequences for Modifying Model Decisions
- Finite-Sample Analysis of Off-Policy Natural Actor-Critic Algorithm
- Intermittent Inference with Nonuniformly Compressed Multi-Exit Neural Network for Energy Harvesting Powered Devices
- Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling
- VFunc: a Deep Generative Model for Functions
- Accelerated Intravascular Ultrasound Imaging using Deep Reinforcement Learning
- Cyrus 2D Simulation Team Description Paper2018
- Data Driven Control with Learned Dynamics: Model-Based versus Model-Free Approach
- Fourier Policy Gradients
- Bayesian Curiosity for Efficient Exploration in Reinforcement Learning
- Deep Reinforcement Learning with Embedded LQR Controllers
- A Multi-Agent Off-Policy Actor-Critic Algorithm for Distributed Reinforcement Learning
- Scalable Synthesis of Verified Controllers in Deep Reinforcement Learning
- QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning
- TAAC: Temporally Abstract Actor-Critic for Continuous Control
- MEPG: A Minimalist Ensemble Policy Gradient Framework for Deep Reinforcement Learning
- A Maximum Mutual Information Framework for Multi-Agent Reinforcement Learning
- Reinforcement Learning vs. Gradient-Based Optimisation for Robust Energy Landscape Control of Spin-1/2 Quantum Networks
- Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models
- Decision Making for Autonomous Driving via Augmented Adversarial Inverse Reinforcement Learning
- O2A: One-shot Observational learning with Action vectors
- Policy Learning Using Weak Supervision
- Reverb: A Framework For Experience Replay
- A Deep Reinforcement Learning Approach for Dynamically Stable Inverse Kinematics of Humanoid Robots
- Continuous Control with Deep Reinforcement Learning for Autonomous Vessels
- Towards continuous control of flippers for a multi-terrain robot using deep reinforcement learning
- Average Reward Adjusted Discounted Reinforcement Learning: Near-Blackwell-Optimal Policies for Real-World Applications
- A Survey of Deep Reinforcement Learning Algorithms for Motion Planning and Control of Autonomous Vehicles
- Model-Based Regularization for Deep Reinforcement Learning with Transcoder Networks
- Deep Reinforcement Learning in Quantitative Algorithmic Trading: A Review
- Hindsight Expectation Maximization for Goal-conditioned Reinforcement Learning
- Online Hyper-parameter Tuning in Off-policy Learning via Evolutionary Strategies
- Simultaneous Energy Harvesting and Information Transmission in a MIMO Full-Duplex System: A Machine Learning-Based Design
- A Regularized Opponent Model with Maximum Entropy Objective
- AI-Driven approach for sustainable extraction of earth's subsurface renewable energy while minimizing seismic activity
- Stimulate the Potential of Robots via Competition
- IRS-Assisted Ambient Backscatter Communications Utilizing Deep Reinforcement Learning
- Colored Noise in PPO: Improved Exploration and Performance through Correlated Action Sampling
- Language Grounding through Social Interactions and Curiosity-Driven Multi-Goal Learning
- Learning Extreme Hummingbird Maneuvers on Flapping Wing Robots
- Balancing Constraints and Rewards with Meta-Gradient D4PG
- Low Emission Building Control with Zero-Shot Reinforcement Learning
- An End-to-end Deep Reinforcement Learning Approach for the Long-term Short-term Planning on the Frenet Space
- Sampled Policy Gradient for Learning to Play the Game Agar.io
- Instabilities of Offline RL with Pre-Trained Neural Representation
- Information Theoretic Model Predictive Q-Learning
- Structured Control Nets for Deep Reinforcement Learning
- Population-Guided Parallel Policy Search for Reinforcement Learning
- Run, skeleton, run: skeletal model in a physics-based simulation
- Control Frequency Adaptation via Action Persistence in Batch Reinforcement Learning
- Reinforcement Learning of Beam Codebooks in Millimeter Wave and Terahertz MIMO Systems
- Towards Lifelong Self-Supervision: A Deep Learning Direction for Robotics
- Improving Computational Efficiency in Visual Reinforcement Learning via Stored Embeddings
- Efficient and Robust Reinforcement Learning with Uncertainty-based Value Expansion
- Practical Reinforcement Learning For MPC: Learning from sparse objectives in under an hour on a real robot
- Locally Persistent Exploration in Continuous Control Tasks with Sparse Rewards
- An Efficient Transfer Learning Framework for Multiagent Reinforcement Learning
- Continuous Doubly Constrained Batch Reinforcement Learning
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement Learning
- Plan2Vec: Unsupervised Representation Learning by Latent Plans
- Hierarchical Approaches for Reinforcement Learning in Parameterized Action Space
- Hierarchical Decomposition of Nonlinear Dynamics and Control for System Identification and Policy Distillation
- How to Close Sim-Real Gap? Transfer with Segmentation!
- Learning Practically Feasible Policies for Online 3D Bin Packing
- PowerNet: Multi-agent Deep Reinforcement Learning for Scalable Powergrid Control
- A Workflow for Offline Model-Free Robotic Reinforcement Learning
- Dynamic Queue-Jump Lane for Emergency Vehicles under Partially Connected Settings: A Multi-Agent Deep Reinforcement Learning Approach
- A Multi-Agent Deep Reinforcement Learning based Spectrum Allocation Framework for D2D Communications
- Implicit Policy for Reinforcement Learning
- Attention-Privileged Reinforcement Learning
- Bridging Imagination and Reality for Model-Based Deep Reinforcement Learning
- Two-stage Deep Reinforcement Learning for Inverter-based Volt-VAR Control in Active Distribution Networks
- Deep Quality-Value (DQV) Learning
- Offline Reinforcement Learning with Pseudometric Learning
- ANT: Learning Accurate Network Throughput for Better Adaptive Video Streaming
- An Actor-Critic-Based UAV-BSs Deployment Method for Dynamic Environments
- Domain Adaptation In Reinforcement Learning Via Latent Unified State Representation
- Intrinsic Motivation for Encouraging Synergistic Behavior
- Breaking the Deadly Triad with a Target Network
- Convergent and Efficient Deep Q Network Algorithm
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-based Approach
- Model-free and Bayesian Ensembling Model-based Deep Reinforcement Learning for Particle Accelerator Control Demonstrated on the FERMI FEL
- POLAR: A Polynomial Arithmetic Framework for Verifying Neural-Network Controlled Systems
- Indoor Point-to-Point Navigation with Deep Reinforcement Learning and Ultra-wideband
- Teach Biped Robots to Walk via Gait Principles and Reinforcement Learning with Adversarial Critics
- Real-time Artificial Intelligence for Accelerator Control: A Study at the Fermilab Booster
- Meta Automatic Curriculum Learning
- Primal-dual Learning for the Model-free Risk-constrained Linear Quadratic Regulator
- Modelling Bounded Rationality in Multi-Agent Interactions by Generalized Recursive Reasoning
- Balance Between Efficient and Effective Learning: Dense2Sparse Reward Shaping for Robot Manipulation with Environment Uncertainty
- Trust Region Value Optimization using Kalman Filtering
- Deep Learning of Koopman Representation for Control
- Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning
- Facilitating Database Tuning with Hyper-Parameter Optimization: A Comprehensive Experimental Evaluation
- Action-Sufficient State Representation Learning for Control with Structural Constraints
- Lyapunov-guided Deep Reinforcement Learning for Stable Online Computation Offloading in Mobile-Edge Computing Networks
- Off-Policy Actor-Critic in an Ensemble: Achieving Maximum General Entropy and Effective Environment Exploration in Deep Reinforcement Learning
- Hashing over Predicted Future Frames for Informed Exploration of Deep Reinforcement Learning
- Generalized Decision Transformer for Offline Hindsight Information Matching
- GRAC: Self-Guided and Self-Regularized Actor-Critic
- Catalyst.RL: A Distributed Framework for Reproducible RL Research
- Adversarial Skill Chaining for Long-Horizon Robot Manipulation via Terminal State Regularization
- FIXAR: A Fixed-Point Deep Reinforcement Learning Platform with Quantization-Aware Training and Adaptive Parallelism
- Symbolic Regression via Neural-Guided Genetic Programming Population Seeding
- Reinforcement Learning Evaluation and Solution for the Feedback Capacity of the Ising Channel with Large Alphabet
- How Much Do Unstated Problem Constraints Limit Deep Robotic Reinforcement Learning?
- Soft Expert Reward Learning for Vision-and-Language Navigation
- Imitation Learning for Human Pose Prediction
- Zeroth-order Deterministic Policy Gradient
- Benchmarking Inverse Optimization Algorithms for Materials Design
- Value-Decomposition Multi-Agent Actor-Critics
- Implicit Distributional Reinforcement Learning
- Is the Policy Gradient a Gradient?
- Automated Lane Change Decision Making using Deep Reinforcement Learning in Dynamic and Uncertain Highway Environment
- Non-local Policy Optimization via Diversity-regularized Collaborative Exploration
- Mutual Information Based Knowledge Transfer Under State-Action Dimension Mismatch
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated Environments
- Generalized Off-Policy Actor-Critic
- Variational Autoencoders for Opponent Modeling in Multi-Agent Systems
- Deep Emotion: A Computational Model of Emotion Using Deep Neural Networks
- Towards Safe Control of Continuum Manipulator Using Shielded Multiagent Reinforcement Learning
- Safe Learning MPC with Limited Model Knowledge and Data
- Mean-Variance Policy Iteration for Risk-Averse Reinforcement Learning
- Balancing Accuracy and Fairness for Interactive Recommendation with Reinforcement Learning
- Reinforcement Learning with Function-Valued Action Spaces for Partial Differential Equation Control
- Deep Learning in Robotics: A Review of Recent Research
- Reinforcement Learning for Quantitative Trading
- Dynamic Control of Stochastic Evolution: A Deep Reinforcement Learning Approach to Adaptively Targeting Emergent Drug Resistance
- Robotic Table Tennis with Model-Free Reinforcement Learning
- AI based Algorithms of Path Planning, Navigation and Control for Mobile Ground Robots and UAVs
- Developing a Simple Model for Sand-Tool Interaction and Autonomously Shaping Sand
- MoTiAC: Multi-Objective Actor-Critics for Real-Time Bidding
- A Visual Communication Map for Multi-Agent Deep Reinforcement Learning
- Incremental Reinforcement Learning --- a New Continuous Reinforcement Learning Frame Based on Stochastic Differential Equation methods
- Learning to Discretize: Solving 1D Scalar Conservation Laws via Deep Reinforcement Learning
- Steady State Analysis of Episodic Reinforcement Learning
- Constrained Model-Free Reinforcement Learning for Process Optimization
- Efficient Deep Reinforcement Learning via Adaptive Policy Transfer
- Deep Reinforcement Learning with Feedback-based Exploration
- Nonparametric Stochastic Compositional Gradient Descent for Q-Learning in Continuous Markov Decision Problems
- Self-supervised Graph Representation Learning via Bootstrapping
- StarCraft Micromanagement with Reinforcement Learning and Curriculum Transfer Learning
- MRAC-RL: A Framework for On-Line Policy Adaptation Under Parametric Model Uncertainty
- Cache-Enabled Dynamic Rate Allocation via Deep Self-Transfer Reinforcement Learning
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation
- Learning to Engage with Interactive Systems: A Field Study on Deep Reinforcement Learning in a Public Museum
- Only Relevant Information Matters: Filtering Out Noisy Samples to Boost RL
- DiGrad: Multi-Task Reinforcement Learning with Shared Actions
- Anti-Jerk On-Ramp Merging Using Deep Reinforcement Learning
- Tailored Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud System
- Reinforcement Learning via Parametric Cost Function Approximation for Multistage Stochastic Programming
- Cooperative Highway Work Zone Merge Control based on Reinforcement Learning in A Connected and Automated Environment
- Wide Area Measurement System-based Low Frequency Oscillation Damping Control through Reinforcement Learning
- Reinforcement Learning with Probabilistically Complete Exploration
- Estimating Q(s,s') with Deep Deterministic Dynamics Gradients
- Deep Reinforcement Learning for Producing Furniture Layout in Indoor Scenes
- UDO: Universal Database Optimization using Reinforcement Learning
- Episodic Linear Quadratic Regulators with Low-rank Transitions
- Iterative Amortized Policy Optimization
- Deep reinforcement learning for the control of conjugate heat transfer with application to workpiece cooling
- Few Shot System Identification for Reinforcement Learning
- Coordination in Adversarial Sequential Team Games via Multi-Agent Deep Reinforcement Learning
- When Do Drivers Concentrate? Attention-based Driver Behavior Modeling With Deep Reinforcement Learning
- UAV-to-Device Underlay Communications: Age of Information Minimization by Multi-agent Deep Reinforcement Learning
- An Off-policy Policy Gradient Theorem Using Emphatic Weightings
- Reward Tweaking: Maximizing the Total Reward While Planning for Short Horizons
- TauRieL: Targeting Traveling Salesman Problem with a deep reinforcement learning inspired architecture
- Curious Exploration and Return-based Memory Restoration for Deep Reinforcement Learning
- Skill Transfer in Deep Reinforcement Learning under Morphological Heterogeneity
- Policy Analysis using Synthetic Controls in Continuous-Time
- Obstacle Avoidance and Navigation Utilizing Reinforcement Learning with Reward Shaping
- The Adversarial Resilience Learning Architecture for AI-based Modelling, Exploration, and Operation of Complex Cyber-Physical Systems
- Multi-Agent Deep Reinforcement Learning based Spectrum Allocation for D2D Underlay Communications
- Evaluating the progress of Deep Reinforcement Learning in the real world: aligning domain-agnostic and domain-specific research
- Integrated Decision and Control: Towards Interpretable and Computationally Efficient Driving Intelligence
- Trajectory Optimization for Unknown Constrained Systems using Reinforcement Learning
- Multiple-objective Reinforcement Learning for Inverse Design and Identification
- Learning Density Distribution of Reachable States for Autonomous Systems
- A User's Guide to Calibrating Robotics Simulators
- Tutoring Reinforcement Learning via Feedback Control
- Soft Policy Gradient Method for Maximum Entropy Deep Reinforcement Learning
- Continuous Homeostatic Reinforcement Learning for Self-Regulated Autonomous Agents
- Data-driven battery operation for energy arbitrage using rainbow deep reinforcement learning
- Learning Navigation Behaviors End-to-End with AutoRL
- Variance-Reduced Off-Policy Memory-Efficient Policy Search
- Intelligent Coordination among Multiple Traffic Intersections Using Multi-Agent Reinforcement Learning
- Dynamic Coded Caching in Wireless Networks Using Multi-Agent Reinforcement Learning
- Multi-Agent Reinforcement Learning for Unmanned Aerial Vehicle Coordination by Multi-Critic Policy Gradient Optimization
- Dampen the Stop-and-Go Traffic with Connected and Automated Vehicles -- A Deep Reinforcement Learning Approach
- Theory-based Causal Transfer: Integrating Instance-level Induction and Abstract-level Structure Learning
- Deep Stock Trading: A Hierarchical Reinforcement Learning Framework for Portfolio Optimization and Order Execution
- Smart Train Operation Algorithms based on Expert Knowledge and Reinforcement Learning
- Synthesizing Chemical Plant Operation Procedures using Knowledge, Dynamic Simulation and Deep Reinforcement Learning
- Learning Multi-agent Skills for Tabular Reinforcement Learning using Factor Graphs
- On the Sample Complexity of Reinforcement Learning with Policy Space Generalization
- Deep Learning with Experience Ranking Convolutional Neural Network for Robot Manipulator
- Self-organization of action hierarchy and compositionality by reinforcement learning with recurrent neural networks
- Deep Reinforcement Learning for Multi-Agent Power Control in Heterogeneous Networks
- Analysing Factorizations of Action-Value Networks for Cooperative Multi-Agent Reinforcement Learning
- Domain-Robust Visual Imitation Learning with Mutual Information Constraints
- Cooperative Internet of UAVs: Distributed Trajectory Design by Multi-agent Deep Reinforcement Learning
- TeachMyAgent: a Benchmark for Automatic Curriculum Learning in Deep RL
- Data-driven Outer-Loop Control Using Deep Reinforcement Learning for Trajectory Tracking
- Stealing Deep Reinforcement Learning Models for Fun and Profit
- Combining Reinforcement Learning with Model Predictive Control for On-Ramp Merging
- Stealthy and Efficient Adversarial Attacks against Deep Reinforcement Learning
- Enhancing a Neurocognitive Shared Visuomotor Model for Object Identification, Localization, and Grasping With Learning From Auxiliary Tasks
- Linear Representation Meta-Reinforcement Learning for Instant Adaptation
- Learning Dynamics Models for Model Predictive Agents
- Low-Bandwidth Communication Emerges Naturally in Multi-Agent Learning Systems
- Deep Reinforcement Learning for Scheduling in Cellular Networks
- Regularized Behavior Value Estimation
- AW-Opt: Learning Robotic Skills with Imitation and Reinforcement at Scale
- Qgraph-bounded Q-learning: Stabilizing Model-Free Off-Policy Deep Reinforcement Learning
- Continuous-Time Model-Based Reinforcement Learning
- Learning Cross-Domain Correspondence for Control with Dynamics Cycle-Consistency
- Sequoia: A Software Framework to Unify Continual Learning Research
- Conformal Bootstrap with Reinforcement Learning
- A Physics-Constrained Deep Learning Model for Simulating Multiphase Flow in 3D Heterogeneous Porous Media
- Caching Transient Content for IoT Sensing: Multi-Agent Soft Actor-Critic
- Multi-Agent Path Planning based on MPC and DDPG
- Flappy Hummingbird: An Open Source Dynamic Simulation of Flapping Wing Robots and Animals
- Solving Challenging Dexterous Manipulation Tasks With Trajectory Optimisation and Reinforcement Learning
- Automated Gain Control Through Deep Reinforcement Learning for Downstream Radar Object Detection
- Robust Constrained Reinforcement Learning for Continuous Control with Model Misspecification
- AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control
- EdgeSlice: Slicing Wireless Edge Computing Network with Decentralized Deep Reinforcement Learning
- Replacing Rewards with Examples: Example-Based Policy Search via Recursive Classification
- Towards Optimal District Heating Temperature Control in China with Deep Reinforcement Learning
- Corner Case Generation and Analysis for Safety Assessment of Autonomous Vehicles
- Reinforcement Learning for Joint Optimization of Multiple Rewards
- Imitation Learning from Pixel-Level Demonstrations by HashReward
- A Learning Framework for High Precision Industrial Assembly
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Policies
- Reciprocal Collision Avoidance for General Nonlinear Agents using Reinforcement Learning
- Value Function Spaces: Skill-Centric State Abstractions for Long-Horizon Reasoning
- Implicit Generative Modeling for Efficient Exploration
- TFPnP: Tuning-free Plug-and-Play Proximal Algorithm with Applications to Inverse Imaging Problems
- Adversarial Intrinsic Motivation for Reinforcement Learning
- DOB-Net: Actively Rejecting Unknown Excessive Time-Varying Disturbances
- Finding Needles in a Moving Haystack: Prioritizing Alerts with Adversarial Reinforcement Learning
- Efficient Neural Interaction Function Search for Collaborative Filtering
- Continuous Control for Searching and Planning with a Learned Model
- Learnable Strategies for Bilateral Agent Negotiation over Multiple Issues
- Multi-Agent Deep Reinforcement Learning with Adaptive Policies
- Towards Return Parity in Markov Decision Processes
- DRAS-CQSim: A Reinforcement Learning based Framework for HPC Cluster Scheduling
- Communication-Efficient Policy Gradient Methods for Distributed Reinforcement Learning
- Competitive Experience Replay
- Correcting Experience Replay for Multi-Agent Communication
- A Crash Course on Reinforcement Learning
- Robust Multi-Agent Reinforcement Learning with Social Empowerment for Coordination and Communication
- CARL: A Benchmark for Contextual and Adaptive Reinforcement Learning
- Reward-estimation variance elimination in sequential decision processes
- Learning Value Functions in Deep Policy Gradients using Residual Variance
- Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery
- Towards Automatic Actor-Critic Solutions to Continuous Control
- PBCS : Efficient Exploration and Exploitation Using a Synergy between Reinforcement Learning and Motion Planning
- Learning to Run with Potential-Based Reward Shaping and Demonstrations from Video Data
- My Body is a Cage: the Role of Morphology in Graph-Based Incompatible Control
- Learning Image-Conditioned Dynamics Models for Control of Under-actuated Legged Millirobots
- Adaptive Synthetic Characters for Military Training
- Q-Networks for Binary Vector Actions
- Belief-Grounded Networks for Accelerated Robot Learning under Partial Observability
- Continuous Deep Hierarchical Reinforcement Learning for Ground-Air Swarm Shepherding
- Nested Reinforcement Learning Based Control for Protective Relays in Power Distribution Systems
- Reinforcement Learning based Proactive Control for Transmission Grid Resilience to Wildfire
- IR-VIC: Unsupervised Discovery of Sub-goals for Transfer in RL
- CIRL: Controllable Imitative Reinforcement Learning for Vision-based Self-driving
- Diversity-based Trajectory and Goal Selection with Hindsight Experience Replay
- Reinforcement Learning Meets Hybrid Zero Dynamics: A Case Study for RABBIT
- HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation
- Mutual-Information Regularization in Markov Decision Processes and Actor-Critic Learning
- Reinforcement Learning for Intelligent Healthcare Systems: A Comprehensive Survey
- Game Theory and Machine Learning in UAVs-Assisted Wireless Communication Networks: A Survey
- Learning to Solve a Rubik's Cube with a Dexterous Hand
- Vision-Dialog Navigation by Exploring Cross-modal Memory
- Proximal Policy Gradient: PPO with Policy Gradient
- Tutorial and Survey on Probabilistic Graphical Model and Variational Inference in Deep Reinforcement Learning
- Backprop-Q: Generalized Backpropagation for Stochastic Computation Graphs
- Efficient Exploration in Constrained Environments with Goal-Oriented Reference Path
- Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
- Experience Augmentation: Boosting and Accelerating Off-Policy Multi-Agent Reinforcement Learning
- Offline Decentralized Multi-Agent Reinforcement Learning
- Hierarchically Decoupled Imitation for Morphological Transfer
- Recruitment-imitation Mechanism for Evolutionary Reinforcement Learning
- On Instrumental Variable Regression for Deep Offline Policy Evaluation
- Composite Q-learning: Multi-scale Q-function Decomposition and Separable Optimization
- Text Generation Based on Generative Adversarial Nets with Latent Variable
- On the Search for Feedback in Reinforcement Learning
- Multiple Access in Dynamic Cell-Free Networks: Outage Performance and Deep Reinforcement Learning-Based Design
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- AutoEG: Automated Experience Grafting for Off-Policy Deep Reinforcement Learning
- From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence Prediction
- ANS: Adaptive Network Scaling for Deep Rectifier Reinforcement Learning Models
- Reinforcement Learning with Efficient Active Feature Acquisition
- Policy Smoothing for Provably Robust Reinforcement Learning
- Competitive Multi-Agent Deep Reinforcement Learning with Counterfactual Thinking
- Learning to Navigate Cloth using Haptics
- Phoebe: Reuse-Aware Online Caching with Reinforcement Learning for Emerging Storage Models
- Auto-Agent-Distiller: Towards Efficient Deep Reinforcement Learning Agents via Neural Architecture Search
- Reinforcement Learning Control of Constrained Dynamic Systems with Uniformly Ultimate Boundedness Stability Guarantee
- CRAVES: Controlling Robotic Arm with a Vision-based Economic System
- Inter-Level Cooperation in Hierarchical Reinforcement Learning
- Evolving Neural Networks in Reinforcement Learning by means of UMDAc
- Pretrain Soft Q-Learning with Imperfect Demonstrations
- Policy Search by Target Distribution Learning for Continuous Control
- Towards Theoretical Understandings of Robust Markov Decision Processes: Sample Complexity and Asymptotics
- Target-Based Temporal Difference Learning
- Robust Reinforcement Learning: A Case Study in Linear Quadratic Regulation
- Using Deep Reinforcement Learning to Learn High-Level Policies on the ATRIAS Biped
- Multi-task Learning with Gradient Guided Policy Specialization
- Task-oriented Design through Deep Reinforcement Learning
- Avoidance of Manual Labeling in Robotic Autonomous Navigation Through Multi-Sensory Semi-Supervised Learning
- Teaching UAVs to Race: End-to-End Regression of Agile Controls in Simulation
- Nano: Nested Human-in-the-Loop Reward Learning for Few-shot Language Model Control
- Supervised Learning and Reinforcement Learning of Feedback Models for Reactive Behaviors: Tactile Feedback Testbed
- Robust Inverse Reinforcement Learning under Transition Dynamics Mismatch
- Bridging the Imitation Gap by Adaptive Insubordination
- Behavior-Guided Actor-Critic: Improving Exploration via Learning Policy Behavior Representation for Deep Reinforcement Learning
- Learning Data Teaching Strategies Via Knowledge Tracing
- Deep Reinforcement Learning with Spatio-temporal Traffic Forecasting for Data-Driven Base Station Sleep Control
- Escaping from Zero Gradient: Revisiting Action-Constrained Reinforcement Learning via Frank-Wolfe Policy Optimization
- Attentional Policies for Cross-Context Multi-Agent Reinforcement Learning
- Kernel-based diffusion approximated Markov decision processes for autonomous navigation and control on unstructured terrains
- Distributed Soft Actor-Critic with Multivariate Reward Representation and Knowledge Distillation
- Learning Agents With Prioritization and Parameter Noise in Continuous State and Action Space
- Aggressive Q-Learning with Ensembles: Achieving Both High Sample Efficiency and High Asymptotic Performance
- Exploration via Hindsight Goal Generation
- Reinforcement Mechanism Design for e-commerce
- Danger-aware Adaptive Composition of DRL Agents for Self-navigation
- A Survey on Reproducibility by Evaluating Deep Reinforcement Learning Algorithms on Real-World Robots
- Efficient Navigation of Colloidal Robots in an Unknown Environment via Deep Reinforcement Learning
- An Empirical Analysis of Proximal Policy Optimization with Kronecker-factored Natural Gradients
- Improved Reinforcement Learning through Imitation Learning Pretraining Towards Image-based Autonomous Driving
- A Dual Memory Structure for Efficient Use of Replay Memory in Deep Reinforcement Learning
- Continual Reinforcement Learning with Diversity Exploration and Adversarial Self-Correction
- Deep Reinforcement Learning with Discrete Normalized Advantage Functions for Resource Management in Network Slicing
- TBQ(): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning
- Fast Skill Learning for Variable Compliance Robotic Assembly
- A* Tree Search for Portfolio Management
- Towards Physically Safe Reinforcement Learning under Supervision
- Active localization of multiple targets using noisy relative measurements
- Learning Pregrasp Manipulation of Objects from Ungraspable Poses
- End-to-End Vision-Based Adaptive Cruise Control (ACC) Using Deep Reinforcement Learning
- Memristor Hardware-Friendly Reinforcement Learning
- Multi-Robot Formation Control Using Reinforcement Learning
- Functional Error Correction for Robust Neural Networks
- Taming an autonomous surface vehicle for path following and collision avoidance using deep reinforcement learning
- Efficient Robotic Task Generalization Using Deep Model Fusion Reinforcement Learning
- Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy Optimization
- A Deep Reinforcement Learning Architecture for Multi-stage Optimal Control
- Adaptive Leader-Follower Formation Control and Obstacle Avoidance via Deep Reinforcement Learning
- Coordination of PV Smart Inverters Using Deep Reinforcement Learning for Grid Voltage Regulation
- Learn to Exceed: Stereo Inverse Reinforcement Learning with Concurrent Policy Optimization
- Competitiveness of MAP-Elites against Proximal Policy Optimization on locomotion tasks in deterministic simulations
- Expressing Diverse Human Driving Behavior with Probabilistic Rewards and Online Inference
- DDPG-based Resource Management for MEC/UAV-Assisted Vehicular Networks
- DeepSlicing: Deep Reinforcement Learning Assisted Resource Allocation for Network Slicing
- Multi-robot Cooperative Object Transportation using Decentralized Deep Reinforcement Learning
- Generative Design of Hardware-aware DNNs
- Simulating multi-exit evacuation using deep reinforcement learning
- Generator and Critic: A Deep Reinforcement Learning Approach for Slate Re-ranking in E-commerce
- Towards Automated Safety Coverage and Testing for Autonomous Vehicles with Reinforcement Learning
- Deep Reinforcement Learning Based Spectrum Allocation in Integrated Access and Backhaul Networks
- Sample Efficient Ensemble Learning with Catalyst.RL
- No-Pain No-Gain: DRL Assisted Optimization in Energy-Constrained CR-NOMA Networks
- Self-Imitation Learning by Planning
- Hierarchical deep reinforcement learning controlled three-dimensional navigation of microrobots in blood vessels
- Multi-Agent Reinforcement Learning of 3D Furniture Layout Simulation in Indoor Graphics Scenes
- Deep Reinforcement Learning for Joint Spectrum and Power Allocation in Cellular Networks
- A Multi-intersection Vehicular Cooperative Control based on End-Edge-Cloud Computing
- Fault-Aware Robust Control via Adversarial Reinforcement Learning
- Detecting Adversarial Patches with Class Conditional Reconstruction Networks
- Learning Trajectories for Visual-Inertial System Calibration via Model-based Heuristic Deep Reinforcement Learning
- Generative Inverse Deep Reinforcement Learning for Online Recommendation
- Smooth Imitation Learning via Smooth Costs and Smooth Policies
- Decentralized Multi-Agent Reinforcement Learning: An Off-Policy Method
- Fully Distributed Actor-Critic Architecture for Multitask Deep Reinforcement Learning
- Feedback Linearization of Car Dynamics for Racing via Reinforcement Learning
- A Simple Approach to Continual Learning by Transferring Skill Parameters
- Improved Soft Actor-Critic: Mixing Prioritized Off-Policy Samples with On-Policy Experience
- A First-Occupancy Representation for Reinforcement Learning
- Flying Through a Narrow Gap Using End-to-end Deep Reinforcement Learning Augmented with Curriculum Learning and Sim2Real
- Scalable Multi-agent Reinforcement Learning Algorithm for Wireless Networks
- An Intelligent Energy Management Framework for Hybrid-Electric Propulsion Systems Using Deep Reinforcement Learning
- On the Robustness of Deep Reinforcement Learning in IRS-Aided Wireless Communications Systems
- Plan-Based Relaxed Reward Shaping for Goal-Directed Tasks
- Embodiment and Computational Creativity
- Boosting Offline Reinforcement Learning with Residual Generative Modeling
- A Deep Reinforcement Learning Approach towards Pendulum Swing-up Problem based on TF-Agents
- Programming and Deployment of Autonomous Swarms using Multi-Agent Reinforcement Learning
- Learning cooperative behaviours in adversarial multi-agent systems
- Energy-Efficient Design for a NOMA assisted STAR-RIS Network with Deep Reinforcement Learning
- Fixed Points in Cyber Space: Rethinking Optimal Evasion Attacks in the Age of AI-NIDS
- Improving Learning from Demonstrations by Learning from Experience
- Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks
- Learning Efficient Multi-Agent Cooperative Visual Exploration
- Continuous Control with Action Quantization from Demonstrations
- Provable Regret Bounds for Deep Online Learning and Control
- Double Deep Q-learning Based Real-Time Optimization Strategy for Microgrids
- Gradient Importance Learning for Incomplete Observations
- Aligning an optical interferometer with beam divergence control and continuous action space
- Towards Deeper Deep Reinforcement Learning with Spectral Normalization
- MapGo: Model-Assisted Policy Optimization for Goal-Oriented Tasks
- Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
- Bias-reduced Multi-step Hindsight Experience Replay for Efficient Multi-goal Reinforcement Learning
- Energy Aware Deep Reinforcement Learning Scheduling for Sensors Correlated in Time and Space
- Bayesian Meta-reinforcement Learning for Traffic Signal Control
- Learning to Represent Action Values as a Hypergraph on the Action Vertices
- Learning Nash Equilibria in Zero-Sum Stochastic Games via Entropy-Regularized Policy Approximation
- Knowledge-Assisted Deep Reinforcement Learning in 5G Scheduler Design: From Theoretical Framework to Implementation
- Energy-based Surprise Minimization for Multi-Agent Value Factorization
- Follow the Object: Curriculum Learning for Manipulation Tasks with Imagined Goals
- Learning Complex Multi-Agent Policies in Presence of an Adversary
- Learning to Sample with Local and Global Contexts in Experience Replay Buffer
- Learning Compositional Neural Programs for Continuous Control
- Distributed Uplink Beamforming in Cell-Free Networks Using Deep Reinforcement Learning
- Zeroth-Order Supervised Policy Improvement
- Neural Lyapunov Redesign
- Novel Policy Seeking with Constrained Optimization
- Towards Cognitive Routing based on Deep Reinforcement Learning
- Learning a generative model for robot control using visual feedback
- Mid-flight Propeller Failure Detection and Control of Propeller-deficient Quadcopter using Reinforcement Learning
- Information-Theoretic Performance Limitations of Feedback Control: Underlying Entropic Laws and Generic Bounds
- Fuzzy Tiling Activations: A Simple Approach to Learning Sparse Representations Online
- Merging Deterministic Policy Gradient Estimations with Varied Bias-Variance Tradeoff for Effective Deep Reinforcement Learning
- Learning Classifiers on Positive and Unlabeled Data with Policy Gradient
- Networked Control of Nonlinear Systems under Partial Observation Using Continuous Deep Q-Learning
- Beyond Exponentially Discounted Sum: Automatic Learning of Return Function
- From semantics to execution: Integrating action planning with reinforcement learning for robotic causal problem-solving
- Baconian: A Unified Open-source Framework for Model-Based Reinforcement Learning
- Personalized Cancer Chemotherapy Schedule: a numerical comparison of performance and robustness in model-based and model-free scheduling methodologies
- Zero-shot Deep Reinforcement Learning Driving Policy Transfer for Autonomous Vehicles based on Robust Control
- Policy Optimization with Model-based Explorations
- ACE: An Actor Ensemble Algorithm for Continuous Control with Tree Search
- Using Deep Reinforcement Learning for the Continuous Control of Robotic Arms
- Switching Isotropic and Directional Exploration with Parameter Space Noise in Deep Reinforcement Learning
- Constrained Exploration and Recovery from Experience Shaping
- Learning Adaptive Display Exposure for Real-Time Advertising
- NavigationNet: A Large-scale Interactive Indoor Navigation Dataset
- Intelligent Trainer for Model-Based Reinforcement Learning
- Un résultat intrigant en commande sans modèle
- Shaping in Practice: Training Wheels to Learn Fast Hopping Directly in Hardware
- Semantics, Representations and Grammars for Deep Learning
- Hierarchical RNNs-Based Transformers MADDPG for Mixed Cooperative-Competitive Environments
- MDP Playground: An Analysis and Debug Testbed for Reinforcement Learning
- MACS: Deep Reinforcement Learning based SDN Controller Synchronization Policy Design
- Task-Based Learning via Task-Oriented Prediction Network with Applications in Finance
- InfoRL: Interpretable Reinforcement Learning using Information Maximization
- Graph Pruning for Model Compression
- Representation of Reinforcement Learning Policies in Reproducing Kernel Hilbert Spaces
- Multi-task Reinforcement Learning with a Planning Quasi-Metric
- TTR-Based Reward for Reinforcement Learning with Implicit Model Priors
- PFPN: Continuous Control of Physically Simulated Characters using Particle Filtering Policy Network
- Deep Reinforcement Learning with Weighted Q-Learning
- Detecting Adversarial Examples in Learning-Enabled Cyber-Physical Systems using Variational Autoencoder for Regression
- Flexible and Efficient Long-Range Planning Through Curious Exploration
- Deep Learning: Our Miraculous Year 1990-1991
- Lachesis: Automatic Partitioning for UDF-Centric Analytics
- Parameter-Based Value Functions
- Decentralized Multi-Agents by Imitation of a Centralized Controller
- Extended Radial Basis Function Controller for Reinforcement Learning
- Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles
- A novel control mode of bionic morphing tail based on deep reinforcement learning
- Learning Accurate Extended-Horizon Predictions of High Dimensional Trajectories
- An Active Learning Framework for Efficient Robust Policy Search
- An Accelerated Fitted Value Iteration Algorithm for MDPs with Finite and Vector-Valued Action Space
- Dream and Search to Control: Latent Space Planning for Continuous Control
- Improving the Exploration of Deep Reinforcement Learning in Continuous Domains using Planning for Policy Search
- Combining Semantic Guidance and Deep Reinforcement Learning For Generating Human Level Paintings
- Towards Neural Knowledge DNA
- Stabilizing Transformer-Based Action Sequence Generation For Q-Learning
- SCAPE: Learning Stiffness Control from Augmented Position Control Experiences
- Safe Learning of Uncertain Environments
- Goal-constrained Sparse Reinforcement Learning for End-to-End Driving
- Don't Forget Your Teacher: A Corrective Reinforcement Learning Framework
- Hybrid Policy Learning for Energy-Latency Tradeoff in MEC-Assisted VR Video Service
- Subgoal-based Reward Shaping to Improve Efficiency in Reinforcement Learning
- Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action Tasks
- Semi-On-Policy Training for Sample Efficient Multi-Agent Policy Gradients
- Context-Based Soft Actor Critic for Environments with Non-stationary Dynamics
- A Survey on Reinforcement Learning-Aided Caching in Mobile Edge Networks
- Game of GANs: Game-Theoretical Models for Generative Adversarial Networks
- Analysis of a Target-Based Actor-Critic Algorithm with Linear Function Approximation
- On the Sample Complexity and Metastability of Heavy-tailed Policy Search in Continuous Control
- Learning Altruistic Behaviours in Reinforcement Learning without External Rewards
- Physics-informed Dyna-Style Model-Based Deep Reinforcement Learning for Dynamic Control
- Tolerance-Guided Policy Learning for Adaptable and Transferrable Delicate Industrial Insertion
- Back to Basics: Deep Reinforcement Learning in Traffic Signal Control
- Learning-based Hamilton-Jacobi-Bellman Methods for Optimal Control
- Offline RL With Resource Constrained Online Deployment
- Off-policy Reinforcement Learning with Optimistic Exploration and Distribution Correction
- Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks
- Learning Pessimism for Robust and Efficient Off-Policy Reinforcement Learning
- A Closer Look at Advantage-Filtered Behavioral Cloning in High-Noise Datasets
- Curriculum Offline Imitation Learning
- Analysis of Model-Free Reinforcement Learning Control Schemes on self-balancing Wheeled Extendible System
- Off-Policy Actor-Critic with Emphatic Weightings
- SEIHAI: A Sample-efficient Hierarchical AI for the MineRL Competition
- Unlocking the Potential of Simulators: Design with RL in Mind
- Continuous Control With Ensemble Deep Deterministic Policy Gradients
- Meta Arcade: A Configurable Environment Suite for Meta-Learning
- Tuning Synaptic Connections instead of Weights by Genetic Algorithm in Spiking Policy Network
- Reliability Quantification of Deep Reinforcement Learning-based Control
- Unsupervised Salient Patch Selection for Data-Efficient Reinforcement Learning
- Worrisome Properties of Neural Network Controllers and Their Symbolic Representations
- A Load Balanced Recommendation Approach
- KnowSR: Knowledge Sharing among Homogeneous Agents in Multi-agent Reinforcement Learning
- Controlling an Inverted Pendulum with Policy Gradient Methods-A Tutorial
- Don't Do What Doesn't Matter: Intrinsic Motivation with Action Usefulness
- Importance Sampling based Exploration in Q Learning
- Low-Dimensional State and Action Representation Learning with MDP Homomorphism Metrics
- HAC Explore: Accelerating Exploration with Hierarchical Reinforcement Learning
- Soft Hierarchical Graph Recurrent Networks for Many-Agent Partially Observable Environments
- ADER:Adapting between Exploration and Robustness for Actor-Critic Methods
- Computation Rate Maximum for Mobile Terminals in UAV-assisted Wireless Powered MEC Networks with Fairness Constraint
- Deictic Image Maps: An Abstraction For Learning Pose Invariant Manipulation Policies
- A Survey of Online Auction Mechanism Design Using Deep Learning Approaches
- SaLinA: Sequential Learning of Agents
- Learning Human Behaviors for Robot-Assisted Dressing
- Policy Optimization with Second-Order Advantage Information
- Coordinated Proximal Policy Optimization
- C-Learning: Horizon-Aware Cumulative Accessibility Estimation
- Deep Reinforcement Learning in Electricity Generation Investment for the Minimization of Long-Term Carbon Emissions and Electricity Costs
- Behavior Planning at Urban Intersections through Hierarchical Reinforcement Learning
- Reinforcement Learning with Time-dependent Goals for Robotic Musicians
- Episodic Self-Imitation Learning with Hindsight
- A survey of benchmarking frameworks for reinforcement learning
- Pareto Deterministic Policy Gradients and Its Application in 5G Massive MIMO Networks
- Towards Recognizing New Semantic Concepts in New Visual Domains
- Optimizing Sponsored Search Ranking Strategy by Deep Reinforcement Learning
- Bootstrapping Motor Skill Learning with Motion Planning
- Instance-Aware Predictive Navigation in Multi-Agent Environments
- Discontinuity-Sensitive Optimal Control Learning by Mixture of Experts
- Comparing Task Simplifications to Learn Closed-Loop Object Picking Using Deep Reinforcement Learning
- Predicting Nanorobot Shapes via Generative Models
- A Hybrid Approach for Reinforcement Learning Using Virtual Policy Gradient for Balancing an Inverted Pendulum
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- Learning image quality assessment by reinforcing task amenable data selection
- Online Policy Gradient for Model Free Learning of Linear Quadratic Regulators with Regret
- Reducing Conservativeness Oriented Offline Reinforcement Learning
- Decentralized Circle Formation Control for Fish-like Robots in the Real-world via Reinforcement Learning
- Learning Sampling Policy for Faster Derivative Free Optimization
- Data-Driven Reinforcement Learning for Virtual Character Animation Control
- Technical Report: Adaptive Control for Linearizable Systems Using On-Policy Reinforcement Learning
- How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents
- Discover the Hidden Attack Path in Multi-domain Cyberspace Based on Reinforcement Learning
- Discrete-to-Deep Supervised Policy Learning
- Reinforced Coloring for End-to-End Instance Segmentation
- Costate-focused models for reinforcement learning
- Extracting Latent State Representations with Linear Dynamics from Rich Observations
- Weakness Analysis of Cyberspace Configuration Based on Reinforcement Learning
- Tomography Based Learning for Load Distribution through Opaque Networks
- Adaptive Energy Management for Real Driving Conditions via Transfer Reinforcement Learning
- Nintendo Super Smash Bros. Melee: An "Untouchable" Agent
- Expanding Motor Skills through Relay Neural Networks
- Self-Adapting Recurrent Models for Object Pushing from Learning in Simulation
- Compare and Select: Video Summarization with Multi-Agent Reinforcement Learning
- Balanced Order Batching with Task-Oriented Graph Clustering
- How does the structure embedded in learning policy affect learning quadruped locomotion?
- SREC: Proactive Self-Remedy of Energy-Constrained UAV-Based Networks via Deep Reinforcement Learning
- Lyapunov-Based Reinforcement Learning for Decentralized Multi-Agent Control
- EDEN: Enabling Energy-Efficient, High-Performance Deep Neural Network Inference Using Approximate DRAM
- Fast-UAP: An Algorithm for Speeding up Universal Adversarial Perturbation Generation with Orientation of Perturbation Vectors
- Learning to Mix n-Step Returns: Generalizing lambda-Returns for Deep Reinforcement Learning
- Off-Policy Policy Gradient Algorithms by Constraining the State Distribution Shift
- Learning to drive via Apprenticeship Learning and Deep Reinforcement Learning
- IMPACT: Importance Weighted Asynchronous Architectures with Clipped Target Networks
- Multi-Issue Bargaining With Deep Reinforcement Learning
- Discrete Action On-Policy Learning with Action-Value Critic
- Genetic-Gated Networks for Deep Reinforcement
- Deep Hierarchical Reinforcement Learning Based Recommendations via Multi-goals Abstraction
- Learning Policies through Quantile Regression
- NeoNav: Improving the Generalization of Visual Navigation via Generating Next Expected Observations
- Deep Reinforcement Learning Based Robot Arm Manipulation with Efficient Training Data through Simulation
- Prioritized Guidance for Efficient Multi-Agent Reinforcement Learning Exploration
- Deep Reinforcement Learning for Personalized Search Story Recommendation
- A Deep Reinforcement Learning Approach to Multi-component Job Scheduling in Edge Computing
- Reinforcement learning with world model
- Large scale continuous-time mean-variance portfolio allocation via reinforcement learning
- Attention-based Deep Reinforcement Learning for Multi-view Environments
- An online evolving framework for advancing reinforcement-learning based automated vehicle control
- Artificial Buildings: Safety, Complexity and a Quantifiable Measure of Beauty
- Maximum Entropy Model-based Reinforcement Learning
- Deep Reinforcement Learning for Inquiry Dialog Policies with Logical Formula Embeddings
- Sim-to-Real Transfer of Robot Learning with Variable Length Inputs
- Functional Regularization for Reinforcement Learning via Learned Fourier Features
- Environment Shaping in Reinforcement Learning using State Abstraction
- Techniques Toward Optimizing Viewability in RTB Ad Campaigns Using Reinforcement Learning
- Deep Reinforcement Learning for Joint Beamwidth and Power Optimization in mmWave Systems
- Reducing the Deployment-Time Inference Control Costs of Deep Reinforcement Learning Agents via an Asymmetric Architecture
- End-to-End Race Driving with Deep Reinforcement Learning
- Deep Reinforcement Learning with Surrogate Agent-Environment Interface
- Micro/Nano Motor Navigation and Localization via Deep Reinforcement Learning
- Learning to Grasp from 2.5D images: a Deep Reinforcement Learning Approach
- Optimizing Intelligent Reflecting Surface-Base Station Association for Mobile Networks
- Variational Policy Search using Sparse Gaussian Process Priors for Learning Multimodal Optimal Actions
- Deep Reinforcement Learning Models Predict Visual Responses in the Brain: A Preliminary Result
- Discrete linear-complexity reinforcement learning in continuous action spaces for Q-learning algorithms
- FiDi-RL: Incorporating Deep Reinforcement Learning with Finite-Difference Policy Search for Efficient Learning of Continuous Control
- Cautious Actor-Critic
- Dimensionality Reduction of Movement Primitives in Parameter Space
- Influence-Based Reinforcement Learning for Intrinsically-Motivated Agents
- Auto Deep Compression by Reinforcement Learning Based Actor-Critic Structure
- Three-Dimensional Trajectory Design for Multi-User MISO UAV Communications: A Deep Reinforcement Learning Approach
- Parallelized Reverse Curriculum Generation
- Learning Task Agnostic Skills with Data-driven Guidance
- Neural Network Repair with Reachability Analysis
- Mixed Reinforcement Learning with Additive Stochastic Uncertainty
- Responsive Regulation of Dynamic UAV Communication Networks Based on Deep Reinforcement Learning
- Communication-Computation Efficient Device-Edge Co-Inference via AutoML
- Eden: A Unified Environment Framework for Booming Reinforcement Learning Algorithms
- Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning
- Hindsight Reward Tweaking via Conditional Deep Reinforcement Learning
- A Logarithmic Barrier Method For Proximal Policy Optimization
- Learning Vision-Guided Dynamic Locomotion Over Challenging Terrains
- Deep Reinforcement Learning for Equal Risk Pricing and Hedging under Dynamic Expectile Risk Measures
- Optimal Network Control in Partially-Controllable Networks
- DSDF: An approach to handle stochastic agents in collaborative multi-agent reinforcement learning
- DROMO: Distributionally Robust Offline Model-based Policy Optimization
- Deep Reinforcement Learning Based Multidimensional Resource Management for Energy Harvesting Cognitive NOMA Communications
- Toward Efficient Federated Learning in Multi-Channeled Mobile Edge Network with Layerd Gradient Compression
- A Model-free Deep Reinforcement Learning Approach To Maneuver A Quadrotor Despite Single Rotor Failure
- Efficiently Training On-Policy Actor-Critic Networks in Robotic Deep Reinforcement Learning with Demonstration-like Sampled Exploration
- Realizing Continual Learning through Modeling a Learning System as a Fiber Bundle
- Learning Neural Parsers with Deterministic Differentiable Imitation Learning
- Multiple-Pilot Collaboration for Advanced Remote Intervention using Reinforcement Learning
- Deep Reinforcement Learning with Adjustments
- Adaptive perturbation adversarial training: based on reinforcement learning
- Biomechanic Posture Stabilisation via Iterative Training of Multi-policy Deep Reinforcement Learning Agents
- Learning Deterministic Policy with Target for Power Control in Wireless Networks
- Locality-Sensitive Experience Replay for Online Recommendation
- Learning Coordinated Tasks using Reinforcement Learning in Humanoids
- An Economy of Neural Networks: Learning from Heterogeneous Experiences
- D2RLIR : an improved and diversified ranking function in interactive recommendation systems based on deep reinforcement learning
- Control of a Nature-inspired Scorpion using Reinforcement Learning
- Models of benthic bipedalism
- Human-Level Control without Server-Grade Hardware
- A Supervised-Learning based Hour-Ahead Demand Response of a Behavior-based HEMS approximating MILP Optimization
- Automatic Goal Generation using Dynamical Distance Learning
- Infer Your Enemies and Know Yourself, Learning in Real-Time Bidding with Partially Observable Opponents
- Continuous Control for High-Dimensional State Spaces: An Interactive Learning Approach
- Deep reinforcement learning for RAN optimization and control
- Soft-Robust Algorithms for Batch Reinforcement Learning
- Dynamic Horizon Value Estimation for Model-based Reinforcement Learning
- A Study of Policy Gradient on a Class of Exactly Solvable Models
- Parameter Critic: a Model Free Variance Reduction Method Through Imperishable Samples
- Cross Learning in Deep Q-Networks
- ACDER: Augmented Curiosity-Driven Experience Replay
- Weighted Entropy Modification for Soft Actor-Critic
- Edge Intelligence for Energy-efficient Computation Offloading and Resource Allocation in 5G Beyond
- Deep Reinforcement Learning and Permissioned Blockchain for Content Caching in Vehicular Edge Computing and Networks
- Towards Learning Controllable Representations of Physical Systems
- Similarity Modeling on Heterogeneous Networks via Automatic Path Discovery
- Self Training Autonomous Driving Agent
- Investigation on the generalization of the Sampled Policy Gradient algorithm
- Multi Pseudo Q-learning Based Deterministic Policy Gradient for Tracking Control of Autonomous Underwater Vehicles
- A Smart Sliding Chinese Pinyin Input Method Editor on Touchscreen
- Reinforcement Learning based Distributed Control of Dissipative Networked Systems
- Suspension Regulation of Medium-low-speed Maglev Trains via Deep Reinforcement Learning
- Conditional Importance Sampling for Off-Policy Learning
- Attention-based Fault-tolerant Approach for Multi-agent Reinforcement Learning Systems
- Convexity and monotonicity in nonlinear optimal control under uncertainty
- Weakly-Supervised Cross-Domain Adaptation for Endoscopic Lesions Segmentation
- Zero-shot Policy Learning with Spatial Temporal RewardDecomposition on Contingency-aware Observation
- State Representation Learning from Demonstration
- MBCAL: Sample Efficient and Variance Reduced Reinforcement Learning for Recommender Systems
- Generative Exploration and Exploitation
- Active Hierarchical Imitation and Reinforcement Learning
- Learning Efficient and Effective Exploration Policies with Counterfactual Meta Policy
- Learning to Prevent Leakage: Privacy-Preserving Inference in the Mobile Cloud
- Learning spatial hearing via innate mechanisms
- Research on Autonomous Maneuvering Decision of UCAV based on Approximate Dynamic Programming
- Long-term Joint Scheduling for Urban Traffic
- When Autonomous Systems Meet Accuracy and Transferability through AI: A Survey
- Continuous Deep Q-Learning with Simulator for Stabilization of Uncertain Discrete-Time Systems
- Accelerating Deep Reinforcement Learning With the Aid of Partial Model: Energy-Efficient Predictive Video Streaming
- CoachNet: An Adversarial Sampling Approach for Reinforcement Learning
- Optimizing Multiple Performance Metrics with Deep GSP Auctions for E-commerce Advertising
- Adjust Planning Strategies to Accommodate Reinforcement Learning Agents
- Learning Robust and Adaptive Real-World Continuous Control Using Simulation and Transfer Learning
- RLINK: Deep Reinforcement Learning for User Identity Linkage
- RACE: Reinforced Cooperative Autonomous Vehicle Collision AvoidancE
- A2: Extracting Cyclic Switchings from DOB-nets for Rejecting Excessive Disturbances
- Randomized Policy Learning for Continuous State and Action MDPs
- Generative Adversarial Imagination for Sample Efficient Deep Reinforcement Learning
- Policy Gradient RL Algorithms as Directed Acyclic Graphs
- Unmanned Surface Vehicle Path Planning from the Perspective of Multi-Modality Constraints: A Comprehensive Analysis
- Continuous Multi-objective Zero-touch Network Slicing via Twin Delayed DDPG and OpenAI Gym
- Deep Learning for Wireless Coded Caching with Unknown and Time-Variant Content Popularity
- Congested Urban Networks Tend to Be Insensitive to Signal Settings: Implications for Learning-Based Control
- Deep Reinforcement Learning using Genetic Algorithm for Parameter Optimization
- Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy
- Efficient Reinforcement Learning Development with RLzoo
- Planning with Exploration: Addressing Dynamics Bottleneck in Model-based Reinforcement Learning
- Generation of Traffic Flows in Multi-Agent Traffic Simulation with Agent Behavior Model based on Deep Reinforcement Learning
- Playing optical tweezers with deep reinforcement learning: in virtual, physical and augmented environments
- Solving Online Threat Screening Games using Constrained Action Space Reinforcement Learning
- Vision-based Robotic Arm Imitation by Human Gesture
- On the Guaranteed Almost Equivalence between Imitation Learning from Observation and Demonstration
- Deep Reinforcement Learning for Backup Strategies against Adversaries
- Evaluating task-agnostic exploration for fixed-batch learning of arbitrary future tasks
- Reinforcement Learning for Beam Pattern Design in Millimeter Wave and Massive MIMO Systems
- Learning How to Solve Bubble Ball
- Reinforcement learning with distance-based incentive/penalty (DIP) updates for highly constrained industrial control systems
- From Persistent Homology to Reinforcement Learning with Applications for Retail Banking
- Reinforcement Learning for Robust Missile Autopilot Design
- Scheduling and Power Control for Wireless Multicast Systems via Deep Reinforcement Learning
- Multi-Step Recurrent Q-Learning for Robotic Velcro Peeling
- Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
- A selected review on reinforcement learning based control for autonomous underwater vehicles
- HILONet: Hierarchical Imitation Learning from Non-Aligned Observations
- How to Train your Quadrotor: A Framework for Consistently Smooth and Responsive Flight Control via Reinforcement Learning
- Meta Reinforcement Learning with Distribution of Exploration Parameters Learned by Evolution Strategies
- Combining Off and On-Policy Training in Model-Based Reinforcement Learning
- Analyzing the Hidden Activations of Deep Policy Networks: Why Representation Matters
- Improved Cooperation by Exploiting a Common Signal
- Simulation Studies on Deep Reinforcement Learning for Building Control with Human Interaction
- Reinforcement Learning for Orientation Estimation Using Inertial Sensors with Performance Guarantee
- Maximum Entropy Reinforcement Learning with Mixture Policies
- Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement Learning
- BlockPuzzle - A Challenge in Physical Reasoning and Generalization for Robot Learning
- Model-free Policy Learning with Reward Gradients
- Model Predictive Actor-Critic: Accelerating Robot Skill Acquisition with Deep Reinforcement Learning
- Dynamic Matching Markets in Power Grid: Concepts and Solution using Deep Reinforcement Learning
- A Deep Deterministic Policy Gradient-based Strategy for Stocks Portfolio Management
- Discriminator Augmented Model-Based Reinforcement Learning
- Hierarchical and Partially Observable Goal-driven Policy Learning with Goals Relational Graph
- GMAC: A Distributional Perspective on Actor-Critic Framework
- Setting Out a Software Stack Capable of Hosting a Virtual ROS-based Competition
- Generative Actor-Critic: An Off-policy Algorithm Using the Push-forward Model
- Combine PPO with NES to Improve Exploration
- Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
- Solving Heterogeneous General Equilibrium Economic Models with Deep Reinforcement Learning
- Time-Aware Q-Networks: Resolving Temporal Irregularity for Deep Reinforcement Learning
- Learning and Exploring Motor Skills with Spacetime Bounds
- Scalable, Decentralized Multi-Agent Reinforcement Learning Methods Inspired by Stigmergy and Ant Colonies
- Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations
- TradeR: Practical Deep Hierarchical Reinforcement Learning for Trade Execution
- Iterative Update and Unified Representation for Multi-Agent Reinforcement Learning
- A Dynamics Perspective of Pursuit-Evasion Games of Intelligent Agents with the Ability to Learn
- Doubly Robust Off-Policy Actor-Critic Algorithms for Reinforcement Learning
- Learning to Reach, Swim, Walk and Fly in One Trial: Data-Driven Control with Scarce Data and Side Information
- Automatic Curricula via Expert Demonstrations
- Deep Deterministic Path Following
- Towards Learning to Play Piano with Dexterous Hands and Touch
- Two-stage training algorithm for AI robot soccer
- Going Beyond Linear RL: Sample Efficient Neural Function Approximation
- Crowdfunding Dynamics Tracking: A Reinforcement Learning Approach
- Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model
- Reinforcement Learning Architectures: SAC, TAC, and ESAC
- A Policy Efficient Reduction Approach to Convex Constrained Deep Reinforcement Learning
- MBDP: A Model-based Approach to Achieve both Robustness and Sample Efficiency via Double Dropout Planning
- Learning to Advertise for Organic Traffic Maximization in E-Commerce Product Feeds
- Deep Reinforcement Learning for Complex Manipulation Tasks with Sparse Feedback
- Direct Random Search for Fine Tuning of Deep Reinforcement Learning Policies
- cube2net: Efficient Query-Specific Network Construction with Data Cube Organization
- Optimizing Quantum Variational Circuits with Deep Reinforcement Learning
- Solving the Real Robot Challenge using Deep Reinforcement Learning
- POAR: Efficient Policy Optimization via Online Abstract State Representation Learning
- Out-of-the-box channel pruned networks
- Deep Reinforcement Learning for Online Control of Stochastic Partial Differential Equations
- Parallel Actors and Learners: A Framework for Generating Scalable RL Implementations
- Is High Variance Unavoidable in RL? A Case Study in Continuous Control
- Training Transition Policies via Distribution Matching for Complex Tasks
- Evolutionary Stochastic Policy Distillation
- -: Adaptive Control with Bayesian Learning
- Universal Policies to Learn Them All
- Unbiased Deep Reinforcement Learning: A General Training Framework for Existing and Future Algorithms
- Deep Reinforcement Learning Based Networked Control with Network Delays for Signal Temporal Logic Specifications
- Continuous Multiagent Control using Collective Behavior Entropy for Large-Scale Home Energy Management
- Hierarchy through Composition with Linearly Solvable Markov Decision Processes
- Model-Free Synthesis via Adversarial Reinforcement Learning
- Computing the Feedback Capacity of Finite State Channels using Reinforcement Learning
- DeCOM: Decomposed Policy for Constrained Cooperative Multi-Agent Reinforcement Learning
- Hybrid BYOL-ViT: Efficient approach to deal with small datasets
- Efficient Use of heuristics for accelerating XCS-based Policy Learning in Markov Games
- Uncertainty-aware Low-Rank Q-Matrix Estimation for Deep Reinforcement Learning
- Calculus of Consent via MARL: Legitimating the Collaborative Governance Supplying Public Goods
- Sequential Dynamic Decision Making with Deep Neural Nets on a Test-Time Budget
- On The Transferability of Deep-Q Networks
- Conditional Neural Architecture Search
- Reinforcement Learning for Volt-Var Control: A Novel Two-stage Progressive Training Strategy
- Explicit Gradient Learning
- SAVER: Safe Learning-Based Controller for Real-Time Voltage Regulation
- Potential Field Guided Actor-Critic Reinforcement Learning
- VisualEnv: visual Gym environments with Blender
- Reinforcement Learning in Topology-based Representation for Human Body Movement with Whole Arm Manipulation
- Model Embedding Model-Based Reinforcement Learning
- Adaptive Experience Selection for Policy Gradient