bias detection 1causal inference 1contextual bandits 1covariate balance 1importance sampling 1kernel methods 1Markov decision processes 1offline learning 1offline reinforcement learning 1off-policy evaluation 1treatment recommendation 1
From the 2 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Evaluating covariate balance for long time horizon Markov decision processes
Joshua Spear, Rebecca Pope, Neil J Sebire
The paper investigates how covariate balance diagnostics can be used to detect hidden confounding and model miss‑specification in offline reinforcement learning for long‑horizon tr…
cs.LG2026
Kernel weighted importance sampling for off-policy evaluation in contextual bandits
Joshua Spear, Matthieu Komorowski, Rebecca Pope +2
The paper introduces Kernel-WIS, a new estimator that uses kernel-weighted importance sampling to evaluate policies offline in contextual bandit settings, offering consistency and…