21 citations · 25 across the 2 of their papers we have counts for
3 papers · 1 filter
Causal Bandits: Learning Good Interventions via Causal Inference
Finnian Lattimore, Tor Lattimore, Mark D. Reid
We study the problem of using causal models to improve the rate at which good interventions can be learned online in a stochastic environment. Our formalism combines multi-arm band…
Compliance-Aware Bandits
Nicolás Della Penna, Mark D. Reid, David Balduzzi
Motivated by clinical trials, we study bandits with observable non-compliance. At each step, the learner chooses an arm, after, instead of observing only the reward, it also observ…
Information, Divergence and Risk for Binary Experiments
Mark D. Reid, Robert C. Williamson
We unify f-divergences, Bregman divergences, surrogate loss bounds (regret bounds), proper scoring rules, matching losses, cost curves, ROC-curves and information. We do this by sy…