collaborators

6 papers

stat.AP2026

Evaluating for the long term: Learnings from industry

Leif Sigerson, Tom Cunningham, Winston Chou +22

Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and shar…

cs.LG2026

Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach

Guang-Yuan Hao, Lars van der Laan, Aurélien Bibaut +1

We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in a new, different environmen…

stat.ML2026

Nonparametric Instrumental Variable Analysis Without Structural Equations: Debiased Inference on Functionals of Inverse Problems with No Solutions

Zikai Shen, Nathan Kallus, Dimitri Meunier +3

We consider debiased inference on finite-dimensional functionals of infinite-dimensional least-squares solutions to inverse problems as a way to avoid having to assume exact soluti…

stat.ML2026

Functional Natural Policy Gradients

Aurelien Bibaut, Houssam Zenati, Thibaud Rahier +1

We propose a cross-fitted debiasing device for policy learning from offline data. A key consequence of the resulting learning principle is regret even for policy classes…

stat.ML2026

Fast Best-in-Class Regret for Contextual Bandits

Samuel Girard, Aurelien Bibaut, Arthur Gretton +2

We study the problem of stochastic contextual bandits in the agnostic setting, where the goal is to compete with the best policy in a given class without assuming realizability or…

stat.ML2026

Efficient Inference after Directionally Stable Adaptive Experiments

Zikai Shen, Houssam Zenati, Nathan Kallus +3

We study inference on scalar-valued pathwise differentiable targets after adaptive data collection, such as a bandit algorithm. We introduce a novel target-specific condition, dire…