Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
Imad Aouali, Victor-Emmanuel Brunel, David Rohde +1
In interactive systems, actions are often correlated, presenting an opportunity for more sample-efficient off-policy evaluation (OPE) and learning (OPL) in large action spaces. We…
cs.LG2024
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
Imad Aouali, Victor-Emmanuel Brunel, David Rohde +1
Off-policy learning (OPL) often involves minimizing a risk estimator based on importance weighting to correct bias from the logging policy used to collect data. However, this metho…