activity
20122022
most citedConsistent On-Line Off-Policy Evaluation

38 citations · 43 across the 6 of their papers we have counts for

collaborators

7 papers

cs.LG2022

SoftTreeMax: Policy Gradient with Tree Search

Gal Dalal, Assaf Hallak, Shie Mannor +1

Policy-gradient methods are widely used for learning control policies. They can be easily distributed to multiple workers and reach state-of-the-art results in many domains. Unfort…

cs.LG20213 cited

On Covariate Shift of Latent Confounders in Imitation and Reinforcement Learning

Guy Tennenholtz, Assaf Hallak, Gal Dalal +3

We consider the problem of using expert data with unobserved confounders for imitation and reinforcement learning. We begin by defining the problem of learning from confounded expe…

stat.ML20172 cited

Automatic Representation for Lifetime Value Recommender Systems

Assaf Hallak, Yishay Mansour, Elad Yom-Tov

Many modern commercial sites employ recommender systems to propose relevant content to users. While most systems are focused on maximizing the immediate gain (clicks, purchases or…

stat.ML201738 cited

Consistent On-Line Off-Policy Evaluation

Assaf Hallak, Shie Mannor

The problem of on-line off-policy evaluation (OPE) has been actively studied in the last decade due to its importance both as a stand-alone problem and as a module in a policy impr…

stat.ML2015

Off-policy evaluation for MDPs with unknown structure

Assaf Hallak, François Schnitzler, Timothy Mann +1

Off-policy learning in dynamic decision problems is essential for providing strong evidence that a new policy is better than the one in use. But how can we prove superiority withou…

stat.ML2015

Contextual Markov Decision Processes

Assaf Hallak, Dotan Di Castro, Shie Mannor

We consider a planning problem where the dynamics and rewards of the environment depend on a hidden static parameter referred to as the context. The objective is to learn a strateg…