1 citations · 1 across the 2 of their papers we have counts for
2 papers
math.PR2020★ 1 cited
The Pendulum Arrangement: Maximizing the Escape Time of Heterogeneous Random Walks
Asaf Cassel, Shie Mannor, Guy Tennenholtz
We identify a fundamental phenomenon of heterogeneous one dimensional random walks: the escape (traversal) time is maximized when the heterogeneity in transition probabilities form…
stat.ML2015
Off-policy evaluation for MDPs with unknown structure
Assaf Hallak, François Schnitzler, Timothy Mann +1
Off-policy learning in dynamic decision problems is essential for providing strong evidence that a new policy is better than the one in use. But how can we prove superiority withou…