1 citations · 1 across the 4 of their papers we have counts for
6 papers · 1 filter
On-Policy and Off-Policy Learning for Large Action Spaces
Imad Aouali
This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main frame…
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
Imad Aouali, Victor-Emmanuel Brunel, David Rohde +1
Off-policy learning (OPL) often involves minimizing a risk estimator based on importance weighting to correct bias from the logging policy used to collect data. However, this metho…
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
Imad Aouali, Victor-Emmanuel Brunel, David Rohde +1
In interactive systems, actions are often correlated, presenting an opportunity for more sample-efficient off-policy evaluation (OPE) and learning (OPL) in large action spaces. We…
Diffusion Models Meet Contextual Bandits
Imad Aouali
Efficient online decision-making in contextual bandits is challenging, as methods without informative priors often suffer from computational or statistical inefficiencies. In this…
Exponential Smoothing for Off-Policy Learning
Imad Aouali, Victor-Emmanuel Brunel, David Rohde +1
Off-policy learning (OPL) aims at finding improved policies from logged bandit data, often by minimizing the inverse propensity scoring (IPS) estimator of the risk. In this work, w…
Combining Reward and Rank Signals for Slate Recommendation
Imad Aouali, Sergey Ivanov, Mike Gartrell +4
We consider the problem of slate recommendation, where the recommender system presents a user with a collection or slate composed of K recommended items at once. If the user finds…