activity
20182024
most citedSafe Exploration in Continuous Action Spaces

275 citations · 341 across the 6 of their papers we have counts for

collaborators

9 papers

cs.LG20223 cited

Optimizing Industrial HVAC Systems with Hierarchical Reinforcement Learning

William Wong, Praneet Dutta, Octavian Voicu +3

Reinforcement learning (RL) techniques have been developed to optimize industrial cooling systems, offering substantial energy savings compared to traditional heuristic policies. A…

cs.LG20223 cited

COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation

Jongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz +4

We consider the offline constrained reinforcement learning (RL) problem, in which the agent aims to compute a policy that maximizes expected return while satisfying given cost cons…

cs.LG2021

Autoregressive Dynamics Models for Offline Policy Evaluation and Optimization

Michael R. Zhang, Tom Le Paine, Ofir Nachum +4

Standard dynamics models for continuous control make use of feedforward computation to predict the conditional distribution of next state and reward given current state and action…

cs.LG202123 cited

Benchmarks for Deep Off-Policy Evaluation

Justin Fu, Mohammad Norouzi, Ofir Nachum +10

Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making. The ability…

cs.LG2020

Robust Constrained Reinforcement Learning for Continuous Control with Model Misspecification

Daniel J. Mankowitz, Dan A. Calian, Rae Jeong +5

Many real-world physical control systems are required to satisfy constraints upon deployment. Furthermore, real-world systems are often subject to effects such as non-stationarity,…

cs.LG202030 cited

Hyperparameter Selection for Offline Reinforcement Learning

Tom Le Paine, Cosmin Paduraru, Andrea Michi +5

Offline reinforcement learning (RL purely from logged data) is an important avenue for deploying RL techniques in real-world scenarios. However, existing hyperparameter selection m…