Soft Actor-Critic With Integer Actions
arXiv:2109.08512
Abstract
Reinforcement learning is well-studied under discrete actions. Integer actions setting is popular in the industry yet still challenging due to its high dimensionality. To this end, we study reinforcement learning under integer actions by incorporating the Soft Actor-Critic (SAC) algorithm with an integer reparameterization. Our key observation for integer actions is that their discrete structure can be simplified using their comparability property. Hence, the proposed integer reparameterization does not need one-hot encoding and is of low dimensionality. Experiments show that the proposed SAC under integer actions is as good as the continuous action version on robot control tasks and outperforms Proximal Policy Optimization on power distribution systems control tasks.
The 2022 American Control Conference (ACC)
References in corpus (6)
- The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
- Soft Actor-Critic for Discrete Action Settings
- Discrete and Continuous Action Representation for Practical RL in Video Games
- Learning to run a Power Network Challenge: a Retrospective Analysis
- PowerGym: A Reinforcement Learning Environment for Volt-Var Control in Power Distribution Systems
- Behavior-Guided Actor-Critic: Improving Exploration via Learning Policy Behavior Representation for Deep Reinforcement Learning