154 citations · 561 across the 27 of their papers we have counts for
Showing 2018 · cs.AIShow all
2 papers · 2 filters
cs.AI2018
Path Consistency Learning in Tsallis Entropy Regularized MDPs
Ofir Nachum, Yinlam Chow, Mohammad Ghavamzadeh
We study the sparse entropy-regularized reinforcement learning (ERL) problem in which the entropy term is a special form of the Tsallis entropy. The optimal policy of this formulat…
cs.AI2018
More Robust Doubly Robust Off-policy Evaluation
Mehrdad Farajtabar, Yinlam Chow, Mohammad Ghavamzadeh
We study the problem of off-policy evaluation (OPE) in reinforcement learning (RL), where the goal is to estimate the performance of a policy from the data generated by another pol…