From the 1 of 6 linked papers with an AI index.
6 papers
On-Policy and Off-Policy Learning for Large Action Spaces
Imad Aouali
The thesis investigates how to learn policies for contextual bandits when the action set is extremely large, covering both on‑policy (interactive) and off‑policy (logged data) sett…
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
Imad Aouali, Otmane Sakhi
Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits. Recent advances in OPL primarily optimize OPE estimato…
Diffusion Models Meet Contextual Bandits
Imad Aouali
Efficient online decision-making in contextual bandits is challenging, as methods without informative priors often suffer from computational or statistical inefficiencies. In this…
Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits
Nicolas Nguyen, Imad Aouali, András György +1
We study the problem of Bayesian fixed-budget best-arm identification (BAI) in structured bandits. We propose an algorithm that uses fixed allocations based on the prior informatio…
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
Imad Aouali, Victor-Emmanuel Brunel, David Rohde +1
In interactive systems, actions are often correlated, presenting an opportunity for more sample-efficient off-policy evaluation (OPE) and learning (OPL) in large action spaces. We…
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
Otmane Sakhi, Imad Aouali, Pierre Alquier +1
This work investigates the offline formulation of the contextual bandit problem, where the goal is to leverage past interactions collected under a behavior policy to evaluate, sele…