activity
20172021
most citedMore Efficient Off-Policy Evaluation through Regularized Targeted Learning

17 citations · 34 across the 6 of their papers we have counts for

collaborators

8 papers

math.ST20214 cited

Sequential causal inference in a single world of connected units

Aurelien Bibaut, Maya Petersen, Nikos Vlassis +2

We consider adaptive designs for a trial involving N individuals that we follow along T time steps. We allow for the variables of one individual to depend on its past and on the pa…

math.ST20201 cited

Sufficient and insufficient conditions for the stochastic convergence of Cesàro means

Aurélien F. Bibaut, Alex Luedtke, Mark J. van der Laan

We study the stochastic convergence of the Cesàro mean of a sequence of random variables. These arise naturally in statistical problems that have a sequential component, where the…

cs.LG20202 cited

Rate-adaptive model selection over a collection of black-box contextual bandit algorithms

Aurélien F. Bibaut, Antoine Chambaz, Mark J. van der Laan

We consider the model selection task in the stochastic contextual bandit setting. Suppose we are given a collection of base contextual bandit algorithms. We provide a master algori…

cs.LG2020

Generalized Policy Elimination: an efficient algorithm for Nonparametric Contextual Bandits

Aurélien F. Bibaut, Antoine Chambaz, Mark J. van der Laan

We propose the Generalized Policy Elimination (GPE) algorithm, an oracle-efficient contextual bandit (CB) algorithm inspired by the Policy Elimination algorithm of \cite{dudik2011}…

cs.LG201917 cited

More Efficient Off-Policy Evaluation through Regularized Targeted Learning

Aurélien F. Bibaut, Ivana Malenica, Nikos Vlassis +1

We study the problem of off-policy evaluation (OPE) in Reinforcement Learning (RL), where the aim is to estimate the performance of a new policy given historical data that may have…

math.ST2019

Fast rates for empirical risk minimization over càdlàg functions with bounded sectional variation norm

Aurélien F. Bibaut, Mark J. van der Laan

Empirical risk minimization over classes functions that are bounded for some version of the variation norm has a long history, starting with Total Variation Denoising (Rudin et al.…