Deep Reinforcement Learning in Large Discrete Action Spaces
arXiv:1512.07679
Abstract
Being able to reason in an environment with a large number of discrete actions is essential to bringing reinforcement learning to a larger class of problems. Recommender systems, industrial plants and language models are only some of the many real-world tasks involving large numbers of discrete actions for which current methods are difficult or even often impossible to apply. An ability to generalize over the set of actions as well as sub-linear complexity relative to the size of the set are both necessary to handle such tasks. Current approaches are not able to provide both of these, which motivates the work in this paper. Our proposed approach leverages prior information about the actions to embed them in a continuous space upon which it can generalize. Additionally, approximate nearest-neighbor methods allow for logarithmic-time lookup complexity relative to the number of actions, which is necessary for time-wise tractable training. This combined approach allows reinforcement learning methods to be applied to large-scale learning problems previously intractable with current methods. We demonstrate our algorithm's abilities on a series of tasks having up to one million actions.
Cited by in corpus (26)
- A Brief Survey of Deep Reinforcement Learning
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- Deep Reinforcement Learning for Online Computation Offloading in Wireless Powered Mobile-Edge Computing Networks
- A Closer Look at Invalid Action Masking in Policy Gradient Algorithms
- Bao: Learning to Steer Query Optimizers
- Learning to Operate an Electric Vehicle Charging Station Considering Vehicle-grid Integration
- Deep reinforcement learning for search, recommendation, and online advertising: a survey
- Jointly Learning to Recommend and Advertise
- Multi-Task Fusion via Reinforcement Learning for Long-Term User Satisfaction in Recommender Systems
- Deep progressive reinforcement learning-based flexible resource scheduling framework for IRS and UAV-assisted MEC system
- Multi-Task Recommendations with Reinforcement Learning
- Model-Free Control for Distributed Stream Data Processing using Deep Reinforcement Learning
- Knowledge-guided Deep Reinforcement Learning for Interactive Recommendation
- Exploration and Regularization of the Latent Action Space in Recommendation
- A Deep Reinforcement Learning-Based Resource Scheduler for Massive MIMO Networks
- An Intelligent Transaction Migration Scheme for RAFT-based Private Blockchain in Internet of Things Applications
- Discrete and continuous representations and processing in deep learning: Looking forward
- Generative Slate Recommendation with Reinforcement Learning
- Quality of service based radar resource management using deep reinforcement learning
- Decentralized Federated Reinforcement Learning for User-Centric Dynamic TFDD Control
- Optimizing Throughput Performance in Distributed MIMO Wi-Fi Networks using Deep Reinforcement Learning
- Joint QoS-Aware Scheduling and Precoding for Massive MIMO Systems via Deep Reinforcement Learning
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons
- Dynamic resource matching in manufacturing using deep reinforcement learning
- AutoAssign+: Automatic Shared Embedding Assignment in Streaming Recommendation
- Jointly-Learned State-Action Embedding for Efficient Reinforcement Learning