bayesian methods 1contextual bandits 1large action spaces 1off-policy learning 1on-policy learning 1
From the 1 of 6 linked papers with an AI index.
Showing stat.MLShow all
3 papers · 1 filter
stat.ML2026
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
Imad Aouali, Otmane Sakhi
Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits. Recent advances in OPL primarily optimize OPE estimato…
stat.ML2025
Prior-Dependent Allocations for Bayesian Fixed-Budget Best-Arm Identification in Structured Bandits
Nicolas Nguyen, Imad Aouali, András György +1
We study the problem of Bayesian fixed-budget best-arm identification (BAI) in structured bandits. We propose an algorithm that uses fixed allocations based on the prior informatio…
stat.ML2024
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
Otmane Sakhi, Imad Aouali, Pierre Alquier +1
This work investigates the offline formulation of the contextual bandit problem, where the goal is to leverage past interactions collected under a behavior policy to evaluate, sele…