A Safe Reinforcement Learning Algorithm for Supervisory Control of Power Plants
arXiv:2401.13020 · doi:10.1016/j.knosys.2024.112312
Abstract
Traditional control theory-based methods require tailored engineering for each system and constant fine-tuning. In power plant control, one often needs to obtain a precise representation of the system dynamics and carefully design the control scheme accordingly. Model-free Reinforcement learning (RL) has emerged as a promising solution for control tasks due to its ability to learn from trial-and-error interactions with the environment. It eliminates the need for explicitly modeling the environment's dynamics, which is potentially inaccurate. However, the direct imposition of state constraints in power plant control raises challenges for standard RL methods. To address this, we propose a chance-constrained RL algorithm based on Proximal Policy Optimization for supervisory control. Our method employs Lagrangian relaxation to convert the constrained optimization problem into an unconstrained objective, where trainable Lagrange multipliers enforce the state constraints. Our approach achieves the smallest distance of violation and violation rate in a load-follow maneuver for an advanced Nuclear Power Plant design.
References in corpus (12)
- Adam: A Method for Stochastic Optimization
- Discovering governing equations from data: Sparse identification of nonlinear dynamical systems
- Training language models to follow instructions with human feedback
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Safe Exploration in Continuous Action Spaces
- Reward Constrained Policy Optimization
- Survey on reinforcement learning for language processing
- Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning
- Design of a Supervisory Control System for Autonomous Operation of Advanced Reactors
- Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning