48 citations · 148 across the 19 of their papers we have counts for
Showing 2021 · cs.LGShow all
2 papers · 2 filters
cs.LG2021★ 2 cited
Conservative and Adaptive Penalty for Model-Based Safe Reinforcement Learning
Yecheng Jason Ma, Andrew Shen, Osbert Bastani +1
Reinforcement Learning (RL) agents in the real world must satisfy safety constraints in addition to maximizing a reward objective. Model-based RL algorithms hold promise for reduci…
cs.LG2021★ 1 cited
State Relevance for Off-Policy Evaluation
Simon P. Shen, Yecheng Jason Ma, Omer Gottesman +1
Importance sampling-based estimators for off-policy evaluation (OPE) are valued for their simplicity, unbiasedness, and reliance on relatively few assumptions. However, the varianc…