2 papers
cs.RO2026
COP-Q: Safety-First Reinforcement Learning for Robot Control via Cholesky-Ordered Projection
Guopeng Li, Moritz A. Zanger, Matthijs T. J. Spaan +1
Safe robot control requires maximizing return while satisfying safety constraints. In off-policy safe reinforcement learning, reward and safety Q-values are commonly learned by sep…
cs.LG2026
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
Guopeng Li, Matthijs T. J. Spaan, Julian F. P. Kooij
When safety is formulated as a limit of cumulative cost, safe reinforcement learning (RL) aims to learn policies that maximize return subject to the cost constraint in data collect…