Dueling Network Architectures for Deep Reinforcement Learning
arXiv:1511.06581
Abstract
In recent years there have been many successes of using deep representations in reinforcement learning. Still, many of these applications use conventional architectures, such as convolutional networks, LSTMs, or auto-encoders. In this paper, we present a new neural network architecture for model-free reinforcement learning. Our dueling network represents two separate estimators: one for the state value function and one for the state-dependent action advantage function. The main benefit of this factoring is to generalize learning across actions without imposing any change to the underlying reinforcement learning algorithm. Our results show that this architecture leads to better policy evaluation in the presence of many similar-valued actions. Moreover, the dueling architecture enables our RL agent to outperform the state-of-the-art on the Atari 2600 domain.
15 pages, 5 figures, and 5 tables
References in corpus (2)
Cited by in corpus (408)
- A Brief Survey of Deep Reinforcement Learning
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- Deep Reinforcement Learning for Multi-Agent Systems: A Review of Challenges, Solutions and Applications
- An Introduction to Deep Reinforcement Learning
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Deep Reinforcement Learning for Cyber Security
- Noisy Networks for Exploration
- Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning
- Deep Reinforcement Learning: An Overview
- Learning a Generic Value-Selection Heuristic Inside a Constraint Programming Solver
- A Review of Deep Reinforcement Learning for Smart Building Energy Management
- Deep Exploration via Bootstrapped DQN
- #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
- Continuous Deep Q-Learning with Model-based Acceleration
- Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning Approach
- Implicit Quantile Networks for Distributional Reinforcement Learning
- A Distributional Perspective on Reinforcement Learning
- A Survey of Deep RL and IL for Autonomous Driving Policy Learning
- Reinforcement Learning for IoT Security: A Comprehensive Survey
- CURL: Contrastive Unsupervised Representations for Reinforcement Learning
- Sample Efficient Actor-Critic with Experience Replay
- Go-Explore: a New Approach for Hard-Exploration Problems
- Medical Image Registration Using Deep Neural Networks: A Comprehensive Review
- Averaged-DQN: Variance Reduction and Stabilization for Deep Reinforcement Learning
- Recent advances in applying deep reinforcement learning for flow control: perspectives and future directions
- QPLEX: Duplex Dueling Multi-Agent Q-Learning
- A Survey of Deep Reinforcement Learning in Video Games
- A Deep Hierarchical Approach to Lifelong Learning in Minecraft
- Towards Monocular Vision based Obstacle Avoidance through Deep Reinforcement Learning
- A Theoretical Analysis of Deep Q-Learning
- Reinforcement Learning in Healthcare: A Survey
- Supervised and Unsupervised Neural Approaches to Text Readability
- IG-RL: Inductive Graph Reinforcement Learning for Massive-Scale Traffic Signal Control
- Reinforcement Learning Algorithms: An Overview and Classification
- RUDDER: Return Decomposition for Delayed Rewards
- Deep Reinforcement Learning and the Deadly Triad
- Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning
- Pseudo-Rehearsal: Achieving Deep Reinforcement Learning without Catastrophic Forgetting
- ViZDoom Competitions: Playing Doom from Pixels
- Accelerated Methods for Deep Reinforcement Learning
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- Deep Reinforcement Learning for Real-Time Optimization of Pumps in Water Distribution Systems
- Learning values across many orders of magnitude
- Controlling an Autonomous Vehicle with Deep Reinforcement Learning
- Ensemble-Based Deep Reinforcement Learning for Chatbots
- Horizon: Facebook's Open Source Applied Reinforcement Learning Platform
- Observe and Look Further: Achieving Consistent Performance on Atari
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks
- Never Give Up: Learning Directed Exploration Strategies
- Towards Cognitive Exploration through Deep Reinforcement Learning for Mobile Robots
- Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning
- ChainerRL: A Deep Reinforcement Learning Library
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- Deep Reinforcement Learning Control of Quantum Cartpoles
- Active One-shot Learning
- Ensemble Quantile Networks: Uncertainty-Aware Reinforcement Learning with Applications in Autonomous Driving
- A Deep Reinforcement Learning Approach for the Meal Delivery Problem
- Towards Characterizing Divergence in Deep Q-Learning
- Adaptive Behavior Generation for Autonomous Driving using Deep Reinforcement Learning with Compact Semantic States
- Deep Hierarchical Reinforcement Learning Algorithm in Partially Observable Markov Decision Processes
- Dynamic Weights in Multi-Objective Deep Reinforcement Learning
- Deep Reinforcement Learning for Robotic Manipulation-The state of the art
- Graying the black box: Understanding DQNs
- Deep reinforcement learning for guidewire navigation in coronary artery phantom
- Local and Global Explanations of Agent Behavior: Integrating Strategy Summaries with Saliency Maps
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- Modelling the Dynamic Joint Policy of Teammates with Attention Multi-agent DDPG
- Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
- Learning End-to-end Multimodal Sensor Policies for Autonomous Navigation
- Reinforcement Learning for Generative AI: State of the Art, Opportunities and Open Research Challenges
- Experience-driven Networking: A Deep Reinforcement Learning based Approach
- Data-Efficient Reinforcement Learning with Self-Predictive Representations
- Dealing with Sparse Rewards in Reinforcement Learning
- Improving Generalization in Reinforcement Learning with Mixture Regularization
- Playing Doom with SLAM-Augmented Deep Reinforcement Learning
- The MineRL 2019 Competition on Sample Efficient Reinforcement Learning using Human Priors
- A Deep Reinforcement Learning Framework for Contention-Based Spectrum Sharing
- Simplified Action Decoder for Deep Multi-Agent Reinforcement Learning
- Almost Optimal Model-Free Reinforcement Learning via Reference-Advantage Decomposition
- Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces
- Value Prediction Network
- A Survey of Deep Learning for Data Caching in Edge Network
- Reward learning from human preferences and demonstrations in Atari
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State Observations
- Factorized Q-Learning for Large-Scale Multi-Agent Systems
- Deep Reinforcement Learning for Efficient Measurement of Quantum Devices
- Rethinking the Implementation Tricks and Monotonicity Constraint in Cooperative Multi-Agent Reinforcement Learning
- Deep Reinforcement Learning with Successor Features for Navigation across Similar Environments
- Prioritized Sequence Experience Replay
- Making Deep Q-learning methods robust to time discretization
- Pretraining Representations for Data-Efficient Reinforcement Learning
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning
- Towards falsifiable interpretability research
- Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder Dimension
- Classification with Costly Features as a Sequential Decision-Making Problem
- Boundary-aware Supervoxel-level Iteratively Refined Interactive 3D Image Segmentation with Multi-agent Reinforcement Learning
- Automated Cloud Provisioning on AWS using Deep Reinforcement Learning
- Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network
- MoËT: Mixture of Expert Trees and its Application to Verifiable Reinforcement Learning
- Deep Reinforcement Learning Methods for Structure-Guided Processing Path Optimization
- Deep reinforcement learning in World-Earth system models to discover sustainable management strategies
- Towards Interpretable Reinforcement Learning Using Attention Augmented Agents
- Hierarchical Imitation and Reinforcement Learning
- An Adaptive Clipping Approach for Proximal Policy Optimization
- Network slicing for vehicular communications: a multi-agent deep reinforcement learning approach
- Safe Option-Critic: Learning Safety in the Option-Critic Architecture
- Learning to Fly via Deep Model-Based Reinforcement Learning
- Deep hierarchical reinforcement agents for automated penetration testing
- Mastering Atari with Discrete World Models
- Revealing systematics in phenomenologically viable flux vacua with reinforcement learning
- Modeling Interactions of Autonomous Vehicles and Pedestrians with Deep Multi-Agent Reinforcement Learning for Collision Avoidance
- Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships
- On Inductive Biases in Deep Reinforcement Learning
- Aggregating E-commerce Search Results from Heterogeneous Sources via Hierarchical Reinforcement Learning
- Generalizable Episodic Memory for Deep Reinforcement Learning
- Analysing Results from AI Benchmarks: Key Indicators and How to Obtain Them
- Revisiting Rainbow: Promoting more Insightful and Inclusive Deep Reinforcement Learning Research
- Sample-Efficient Deep Reinforcement Learning via Episodic Backward Update
- Exploration-Enhanced POLITEX
- Learning Robust Dialog Policies in Noisy Environments
- A Human Mixed Strategy Approach to Deep Reinforcement Learning
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- Applications of Deep Reinforcement Learning in Communications and Networking: A Survey
- Building Safer Autonomous Agents by Leveraging Risky Driving Behavior Knowledge
- The Wireless Control Plane: An Overview and Directions for Future Research
- An Information-Theoretic Optimality Principle for Deep Reinforcement Learning
- DRLDO: A novel DRL based De-ObfuscationSystem for Defense against Metamorphic Malware
- Reinforcement Learning with Quantum Variational Circuits
- A Finite-Time Analysis of Q-Learning with Neural Network Function Approximation
- Robust Deep Reinforcement Learning through Adversarial Loss
- Deep Reinforcement Learning for Task Offloading in Mobile Edge Computing Systems
- AI-Based Autonomous Line Flow Control via Topology Adjustment for Maximizing Time-Series ATCs
- Constructions in combinatorics via neural networks
- Hindsight Goal Ranking on Replay Buffer for Sparse Reward Environment
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data Augmentation
- Modern Deep Reinforcement Learning Algorithms
- Reinforcement Learning with Perturbed Rewards
- Learning to Navigate in Indoor Environments: from Memorizing to Reasoning
- Multi-agent Reinforcement Learning Accelerated MCMC on Multiscale Inversion Problem
- Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset
- Recent Advances in Leveraging Human Guidance for Sequential Decision-Making Tasks
- FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling
- The Advantage of Doubling: A Deep Reinforcement Learning Approach to Studying the Double Team in the NBA
- Deep Reinforcement Learning for Autonomous Internet of Things: Model, Applications and Challenges
- Transparency and Explanation in Deep Reinforcement Learning Neural Networks
- Deep Reinforcement Learning for Green Security Games with Real-Time Information
- Applications of Multi-Agent Reinforcement Learning in Future Internet: A Comprehensive Survey
- Ctrl-Z: Recovering from Instability in Reinforcement Learning
- Learn to Interpret Atari Agents
- LADDER: A Human-Level Bidding Agent for Large-Scale Real-Time Online Auctions
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Online 3D Bin Packing with Constrained Deep Reinforcement Learning
- On the Estimation Bias in Double Q-Learning
- Learning to run a Power Network Challenge: a Retrospective Analysis
- Handover Control in Wireless Systems via Asynchronous Multi-User Deep Reinforcement Learning
- World Discovery Models
- Online Robustness Training for Deep Reinforcement Learning
- Spectral Normalisation for Deep Reinforcement Learning: an Optimisation Perspective
- Beyond Fine-Tuning: Transferring Behavior in Reinforcement Learning
- Investigation of Error Simulation Techniques for Learning Dialog Policies for Conversational Error Recovery
- Deep Reinforcement Learning for Trading
- Learning Heuristic Search via Imitation
- Action Semantics Network: Considering the Effects of Actions in Multiagent Systems
- Deep Learning based Wireless Resource Allocation with Application to Vehicular Networks
- Internet of Things Meets Brain-Computer Interface: A Unified Deep Learning Framework for Enabling Human-Thing Cognitive Interactivity
- Double Prioritized State Recycled Experience Replay
- Student-Initiated Action Advising via Advice Novelty
- GAN-powered Deep Distributional Reinforcement Learning for Resource Management in Network Slicing
- DREAM: Deep Regret minimization with Advantage baselines and Model-free learning
- Faster and Safer Training by Embedding High-Level Knowledge into Deep Reinforcement Learning
- Using a Logarithmic Mapping to Enable Lower Discount Factors in Reinforcement Learning
- Improving Robustness of Reinforcement Learning for Power System Control with Adversarial Training
- Actor-Critic Reinforcement Learning for Control with Stability Guarantee
- Deep Reinforcement Learning Based High-level Driving Behavior Decision-making Model in Heterogeneous Traffic
- Divergence-Augmented Policy Optimization
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience Replay
- Simultaneous Navigation and Radio Mapping for Cellular-Connected UAV with Deep Reinforcement Learning
- General non-linear Bellman equations
- Learning-Driven Exploration for Reinforcement Learning
- Adaptive Trade-Offs in Off-Policy Learning
- Using RGB Image as Visual Input for Mapless Robot Navigation
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
- Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space
- Reinforcement Learning-based Switching Controller for a Milliscale Robot in a Constrained Environment
- TD or not TD: Analyzing the Role of Temporal Differencing in Deep Reinforcement Learning
- The Benchmark Lottery
- Text as Environment: A Deep Reinforcement Learning Text Readability Assessment Model
- Modeling 3D Shapes by Reinforcement Learning
- PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement Learning
- Automated Image Data Preprocessing with Deep Reinforcement Learning
- Direct and indirect reinforcement learning
- Variational Deep Q Network
- The PlayStation Reinforcement Learning Environment (PSXLE)
- Self-Imitation Advantage Learning
- Improving Generalization of Reinforcement Learning with Minimax Distributional Soft Actor-Critic
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning without Sacrifices
- Robot Navigation with Map-Based Deep Reinforcement Learning
- Know Your Mind: Adaptive Brain Signal Classification with Reinforced Attentive Convolutional Neural Networks
- Return-based Scaling: Yet Another Normalisation Trick for Deep RL
- Carl-Lead: Lidar-based End-to-End Autonomous Driving with Contrastive Deep Reinforcement Learning
- Efficient Per-Example Gradient Computations in Convolutional Neural Networks
- AgentGraph: Towards Universal Dialogue Management with Structured Deep Reinforcement Learning
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement Learning
- Reinforcement Learning-Based Coverage Path Planning with Implicit Cellular Decomposition
- A Hybrid Stochastic Policy Gradient Algorithm for Reinforcement Learning
- Deep Reinforcement Learning for Demand Driven Services in Logistics and Transportation Systems: A Survey
- On the Reduction of Variance and Overestimation of Deep Q-Learning
- Learning Sampling Policies for Domain Adaptation
- Complementary reinforcement learning towards explainable agents
- Flatland Competition 2020: MAPF and MARL for Efficient Train Coordination on a Grid World
- Dynamic Measurement Scheduling for Event Forecasting using Deep RL
- Distributional Reinforcement Learning with Unconstrained Monotonic Neural Networks
- Distributed Heuristic Multi-Agent Path Finding with Communication
- Structured Control Nets for Deep Reinforcement Learning
- A Reinforcement Learning Approach for Intelligent Traffic Signal Control at Urban Intersections
- Learning Index Selection with Structured Action Spaces
- Policy Learning Using Weak Supervision
- Continuous Doubly Constrained Batch Reinforcement Learning
- A Visual Communication Map for Multi-Agent Deep Reinforcement Learning
- Hierarchical Reinforcement Learning for Relay Selection and Power Optimization in Two-Hop Cooperative Relay Network
- Automated Lane Change Decision Making using Deep Reinforcement Learning in Dynamic and Uncertain Highway Environment
- InFlow: Robust outlier detection utilizing Normalizing Flows
- Convergent and Efficient Deep Q Network Algorithm
- Auto-MAP: A DQN Framework for Exploring Distributed Execution Plans for DNN Workloads
- Challenges of Context and Time in Reinforcement Learning: Introducing Space Fortress as a Benchmark
- StarCraft Micromanagement with Reinforcement Learning and Curriculum Transfer Learning
- Scalable and Incremental Learning of Gaussian Mixture Models
- A Survey of Exploration Methods in Reinforcement Learning
- Bi-level Off-policy Reinforcement Learning for Volt/VAR Control Involving Continuous and Discrete Devices
- Approximating meta-heuristics with homotopic recurrent neural networks
- V2I Connectivity-Based Dynamic Queue-Jump Lane for Emergency Vehicles: A Deep Reinforcement Learning Approach
- Gamifying the Vehicle Routing Problem with Stochastic Requests
- Reinforcement Learning for Joint Optimization of Multiple Rewards
- Learning to Search in Long Documents Using Document Structure
- Deep Reinforcement Learning for Optimal Stopping with Application in Financial Engineering
- Evolutionarily-Curated Curriculum Learning for Deep Reinforcement Learning Agents
- Comprehensive and Efficient Data Labeling via Adaptive Model Scheduling
- Decoupling Value and Policy for Generalization in Reinforcement Learning
- POMDPs in Continuous Time and Discrete Spaces
- Weighing Counts: Sequential Crowd Counting by Reinforcement Learning
- Driving Tasks Transfer in Deep Reinforcement Learning for Decision-making of Autonomous Vehicles
- Variance Reduction for Deep Q-Learning using Stochastic Recursive Gradient
- Robust Image Matching By Dynamic Feature Selection
- Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
- Taylor Expansion Policy Optimization
- Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems
- Convergence Analysis of No-Regret Bidding Algorithms in Repeated Auctions
- Adaptive Traffic Control with Deep Reinforcement Learning: Towards State-of-the-art and Beyond
- Interactive Recommender System via Knowledge Graph-enhanced Reinforcement Learning
- Automatic Data Augmentation by Learning the Deterministic Policy
- Representation Learning of Pedestrian Trajectories Using Actor-Critic Sequence-to-Sequence Autoencoder
- Efficient Reinforcement Learning for StarCraft by Abstract Forward Models and Transfer Learning
- ANS: Adaptive Network Scaling for Deep Rectifier Reinforcement Learning Models
- GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning
- Optimization-Based Algebraic Multigrid Coarsening Using Reinforcement Learning
- An End-to-End Deep RL Framework for Task Arrangement in Crowdsourcing Platforms
- Relationship Explainable Multi-objective Reinforcement Learning with Semantic Explainability Generation
- Target-Based Temporal Difference Learning
- Trajectory and Passive Beamforming Design for IRS-aided Multi-Robot NOMA Indoor Networks
- Danger-aware Adaptive Composition of DRL Agents for Self-navigation
- AdaDeep: A Usage-Driven, Automated Deep Model Compression Framework for Enabling Ubiquitous Intelligent Mobiles
- Program Synthesis Through Reinforcement Learning Guided Tree Search
- Deep Reinforcement Learning with Quantum-inspired Experience Replay
- Reinforcement Learning for Flexibility Design Problems
- Learning Synthetic Environments for Reinforcement Learning with Evolution Strategies
- Zero-Shot Learning of Text Adventure Games with Sentence-Level Semantics
- Finding Needles in a Moving Haystack: Prioritizing Alerts with Adversarial Reinforcement Learning
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- Domain-invariant NBV Planner for Active Cross-domain Self-localization
- Co-training for Policy Learning
- Bayesian Optimization for Iterative Learning
- Relationship Explainable Multi-objective Optimization Via Vector Value Function Based Reinforcement Learning
- Mixture of Step Returns in Bootstrapped DQN
- Learning Abstract Models for Strategic Exploration and Fast Reward Transfer
- Auto-Agent-Distiller: Towards Efficient Deep Reinforcement Learning Agents via Neural Architecture Search
- Generalization Tower Network: A Novel Deep Neural Network Architecture for Multi-Task Learning
- Deep Reinforcement Learning Based Dynamic Route Planning for Minimizing Travel Time
- Deep Reinforcement Learning for Backscatter-Aided Data Offloading in Mobile Edge Computing
- Approximating two value functions instead of one: towards characterizing a new family of Deep Reinforcement Learning algorithms
- Reinforcement Learning with Latent Flow
- Diluted Near-Optimal Expert Demonstrations for Guiding Dialogue Stochastic Policy Optimisation
- Distributed Deep Reinforcement Learning: An Overview
- Temporal-Difference Value Estimation via Uncertainty-Guided Soft Updates
- Improved Soft Actor-Critic: Mixing Prioritized Off-Policy Samples with On-Policy Experience
- Autonomous Curiosity for Real-Time Training Onboard Robotic Agents
- Explainable Deep Reinforcement Learning Using Introspection in a Non-episodic Task
- Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings
- Unbiased Methods for Multi-Goal Reinforcement Learning
- Optimal Status Update for Caching Enabled IoT Networks: A Dueling Deep R-Network Approach
- Reinforced Few-Shot Acquisition Function Learning for Bayesian Optimization
- Deep Reinforcement Learning Based Spectrum Allocation in Integrated Access and Backhaul Networks
- Temporally-Extended ε-Greedy Exploration
- Forgetful Experience Replay in Hierarchical Reinforcement Learning from Demonstrations
- Direct Advantage Estimation
- Expert Q-learning: Deep Reinforcement Learning with Coarse State Values from Offline Expert Examples
- Towards Deeper Deep Reinforcement Learning with Spectral Normalization
- Group Equivariant Deep Reinforcement Learning
- Off-Belief Learning
- Solving optimal stopping problems with Deep Q-Learning
- Revocable Deep Reinforcement Learning with Affinity Regularization for Outlier-Robust Graph Matching
- Simulating multi-exit evacuation using deep reinforcement learning
- Learning to Represent Action Values as a Hypergraph on the Action Vertices
- Dueling Deep Q Network for Highway Decision Making in Autonomous Vehicles: A Case Study
- Energy-based Surprise Minimization for Multi-Agent Value Factorization
- A Comparative Analysis of Deep Reinforcement Learning-enabled Freeway Decision-making for Automated Vehicles
- Human and Multi-Agent collaboration in a human-MARL teaming framework
- EEG-based Drowsiness Estimation for Driving Safety using Deep Q-Learning
- Sample-Efficient Reinforcement Learning with Maximum Entropy Mellowmax Episodic Control
- Memory-Efficient Episodic Control Reinforcement Learning with Dynamic Online k-means
- Improved robustness of reinforcement learning policies upon conversion to spiking neuronal network platforms applied to ATARI games
- Dynamic Measurement Scheduling for Adverse Event Forecasting using Deep RL
- Informative Path Planning for Mobile Sensing with Reinforcement Learning
- Continual Reinforcement Learning with Diversity Exploration and Adversarial Self-Correction
- Discovering an Aid Policy to Minimize Student Evasion Using Offline Reinforcement Learning
- Index Selection for NoSQL Database with Deep Reinforcement Learning
- Tuning Synaptic Connections instead of Weights by Genetic Algorithm in Spiking Policy Network
- Task-Relevant Object Discovery and Categorization for Playing First-person Shooter Games
- Exploiting Multiple Abstractions in Episodic RL via Reward Shaping
- ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents
- Generative Adversarial Exploration for Reinforcement Learning
- Applications of Artificial Intelligence, Machine Learning and related techniques for Computer Networking Systems
- Energy-Efficient Parking Analytics System using Deep Reinforcement Learning
- VRGym: A Virtual Testbed for Physical and Interactive AI
- Deictic Image Maps: An Abstraction For Learning Pose Invariant Manipulation Policies
- Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement Learning
- An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning
- InferNet for Delayed Reinforcement Tasks: Addressing the Temporal Credit Assignment Problem
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy Evaluation
- Constrained Policy Improvement for Safe and Efficient Reinforcement Learning
- Deep Reinforcement Learning with Weighted Q-Learning
- Quality-Aware Multimodal Saliency Detection via Deep Reinforcement Learning
- Adversary Agnostic Robust Deep Reinforcement Learning
- Towards robust and domain agnostic reinforcement learning competitions
- Learning Visual Affordances with Target-Orientated Deep Q-Network to Grasp Objects by Harnessing Environmental Fixtures
- ConQUR: Mitigating Delusional Bias in Deep Q-learning
- Design of Artificial Intelligence Agents for Games using Deep Reinforcement Learning
- Ranking Policy Decisions
- A Novel Deep Reinforcement Learning Based Stock Direction Prediction using Knowledge Graph and Community Aware Sentiments
- Prioritized Guidance for Efficient Multi-Agent Reinforcement Learning Exploration
- Visual Explanation using Attention Mechanism in Actor-Critic-based Deep Reinforcement Learning
- Clustered Reinforcement Learning
- Independent Reinforcement Learning for Weakly Cooperative Multiagent Traffic Control Problem
- Multi-Agent Path Planning Using Deep Reinforcement Learning
- A Survey on Reinforcement Learning-Aided Caching in Mobile Edge Networks
- Classification with Costly Features in Hierarchical Deep Sets
- Ensemble Bootstrapping for Q-Learning
- Optimized Recommender Systems with Deep Reinforcement Learning
- Physics-informed Dyna-Style Model-Based Deep Reinforcement Learning for Dynamic Control
- Comparing Heuristics, Constraint Optimization, and Reinforcement Learning for an Industrial 2D Packing Problem
- Decentralized Multi-Agents by Imitation of a Centralized Controller
- Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks
- Implicitly Regularized RL with Implicit Q-Values
- Reinforcement Learning and Video Games
- Scalable Deep Reinforcement Learning for Routing and Spectrum Access in Physical Layer
- MACS: Deep Reinforcement Learning based SDN Controller Synchronization Policy Design
- Distributed Deep Reinforcement Learning for Collaborative Spectrum Sharing
- Deep Reinforcement Learning for Personalized Search Story Recommendation
- Tactical Reward Shaping: Bypassing Reinforcement Learning with Strategy-Based Goals
- Individual specialization in multi-task environments with multiagent reinforcement learners
- Improving Experience Replay through Modeling of Similar Transitions' Sets
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- Zero-Shot Adaptation for mmWave Beam-Tracking on Overhead Messenger Wires through Robust Adversarial Reinforcement Learning
- Which Channel to Ask My Question? Personalized Customer Service Request Stream Routing using Deep Reinforcement Learning
- Improving On-policy Learning with Statistical Reward Accumulation
- Shared Learning : Enhancing Reinforcement in -Ensembles
- MBCAL: Sample Efficient and Variance Reduced Reinforcement Learning for Recommender Systems
- Application of Deep Q-Network in Portfolio Management
- Learning Transferable Concepts in Deep Reinforcement Learning
- Ranking Policy Gradient
- DSP: A Differential Spatial Prediction Scheme for Comprehensive real industrial datasets
- An adaptive synchronization approach for weights of deep reinforcement learning
- Deep Reinforcement Learning for Tactile Robotics: Learning to Type on a Braille Keyboard
- Efficient Reinforcement Learning Development with RLzoo
- Test-Cost Sensitive Methods for Identifying Nearby Points
- Relational Mimic for Visual Adversarial Imitation Learning
- Cognitive Radio Network Throughput Maximization with Deep Reinforcement Learning
- Agent with Warm Start and Adaptive Dynamic Termination for Plane Localization in 3D Ultrasound
- Time-Aware Q-Networks: Resolving Temporal Irregularity for Deep Reinforcement Learning
- Simultaneous Navigation and Construction Benchmarking Environments
- Dynamic Multichannel Access via Multi-agent Reinforcement Learning: Throughput and Fairness Guarantees
- Autonomous quadrotor obstacle avoidance based on dueling double deep recurrent Q-learning with monocular vision
- Adversarial Reinforcement Learning in Dynamic Channel Access and Power Control
- Did I do that? Blame as a means to identify controlled effects in reinforcement learning
- When should agents explore?
- Object-sensitive Deep Reinforcement Learning
- A Policy Efficient Reduction Approach to Convex Constrained Deep Reinforcement Learning
- CubeTR: Learning to Solve The Rubiks Cube Using Transformers
- On The Transferability of Deep-Q Networks
- Count-Based Temperature Scheduling for Maximum Entropy Reinforcement Learning
- Learning to Learn in Simulation
- Evolution of Q Values for Deep Q Learning in Stable Baselines
- Hi-Phy: A Benchmark for Hierarchical Physical Reasoning
- Interpretable UAV Collision Avoidance using Deep Reinforcement Learning
- Deep Reinforcement Learning using Genetic Algorithm for Parameter Optimization
- Robust Dual View Deep Agent
- Modelling resource allocation in uncertain system environment through deep reinforcement learning
- Deep Reinforcement Learning Models Predict Visual Responses in the Brain: A Preliminary Result
- Building Intelligent Autonomous Navigation Agents
- High Performance Across Two Atari Paddle Games Using the Same Perceptual Control Architecture Without Training
- An Exploration of Deep Learning Methods in Hungry Geese
- Two-stage training algorithm for AI robot soccer
- In Hindsight: A Smooth Reward for Steady Exploration
- Theoretically Principled Deep RL Acceleration via Nearest Neighbor Function Approximation
- Leveraging the Variance of Return Sequences for Exploration Policy
- Continuous Deep Q-Learning with Simulator for Stabilization of Uncertain Discrete-Time Systems
- Task-Oriented Language Grounding for Language Input with Multiple Sub-Goals of Non-Linear Order
- libGroomRL: Reinforcement Learning for Jets