38 citations · 43 across the 6 of their papers we have counts for
7 papers
SoftTreeMax: Policy Gradient with Tree Search
Gal Dalal, Assaf Hallak, Shie Mannor +1
Policy-gradient methods are widely used for learning control policies. They can be easily distributed to multiple workers and reach state-of-the-art results in many domains. Unfort…
On Covariate Shift of Latent Confounders in Imitation and Reinforcement Learning
Guy Tennenholtz, Assaf Hallak, Gal Dalal +3
We consider the problem of using expert data with unobserved confounders for imitation and reinforcement learning. We begin by defining the problem of learning from confounded expe…
Automatic Representation for Lifetime Value Recommender Systems
Assaf Hallak, Yishay Mansour, Elad Yom-Tov
Many modern commercial sites employ recommender systems to propose relevant content to users. While most systems are focused on maximizing the immediate gain (clicks, purchases or…
Consistent On-Line Off-Policy Evaluation
Assaf Hallak, Shie Mannor
The problem of on-line off-policy evaluation (OPE) has been actively studied in the last decade due to its importance both as a stand-alone problem and as a module in a policy impr…
Off-policy evaluation for MDPs with unknown structure
Assaf Hallak, François Schnitzler, Timothy Mann +1
Off-policy learning in dynamic decision problems is essential for providing strong evidence that a new policy is better than the one in use. But how can we prove superiority withou…
Contextual Markov Decision Processes
Assaf Hallak, Dotan Di Castro, Shie Mannor
We consider a planning problem where the dynamics and rewards of the environment depend on a hidden static parameter referred to as the context. The objective is to learn a strateg…