5 citations · 5 across the 2 of their papers we have counts for
3 papers
cs.LG2022
ARMOR: A Model-based Framework for Improving Arbitrary Baseline Policies with Offline Data
Tengyang Xie, Mohak Bhardwaj, Nan Jiang +1
We propose a new model-based offline RL framework, called Adversarial Models for Offline Reinforcement Learning (ARMOR), which can robustly learn policies to improve upon an arbitr…
cs.LG2021★ 5 cited
Safe Reinforcement Learning Using Advantage-Based Intervention
Nolan Wagener, Byron Boots, Ching-An Cheng
Many sequential decision problems involve finding a policy that maximizes total reward while obeying safety constraints. Although much recent research has focused on the developmen…
cs.LG2021
Cautiously Optimistic Policy Optimization and Exploration with Linear Function Approximation
Andrea Zanette, Ching-An Cheng, Alekh Agarwal
Policy optimization methods are popular reinforcement learning algorithms, because their incremental and on-policy nature makes them more stable than the value-based counterparts.…