works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.LG2026

On-Policy and Off-Policy Learning for Large Action Spaces

Imad Aouali

The thesis investigates how to learn policies for contextual bandits when the action set is extremely large, covering both on‑policy (interactive) and off‑policy (logged data) sett…

stat.ML2026

Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation

Imad Aouali, Otmane Sakhi

Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits. Recent advances in OPL primarily optimize OPE estimato…

cs.LG2025

Diffusion Models Meet Contextual Bandits

Imad Aouali

Efficient online decision-making in contextual bandits is challenging, as methods without informative priors often suffer from computational or statistical inefficiencies. In this…

stat.ML2025

Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits

Nicolas Nguyen, Imad Aouali, András György +1

We study the problem of Bayesian fixed-budget best-arm identification (BAI) in structured bandits. We propose an algorithm that uses fixed allocations based on the prior informatio…

cs.LG2025

Bayesian Off-Policy Evaluation and Learning for Large Action Spaces

Imad Aouali, Victor-Emmanuel Brunel, David Rohde +1

In interactive systems, actions are often correlated, presenting an opportunity for more sample-efficient off-policy evaluation (OPE) and learning (OPL) in large action spaces. We…

stat.ML2024

Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning

Otmane Sakhi, Imad Aouali, Pierre Alquier +1

This work investigates the offline formulation of the contextual bandit problem, where the goal is to leverage past interactions collected under a behavior policy to evaluate, sele…