2 papers
cs.LG2026
Offline-to-Online Learning in Linear Bandits
Kushagra Chandak, Toshinori Kitamura, Xiaoqi Tan
We study online learning with an additional offline dataset in the stochastic linear bandit setting. Although this problem arises frequently in practice, the offline-to-online trad…
cs.LG2025
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
Kushagra Chandak, Vincent Liu, Haanvid Lee
We consider off-policy evaluation (OPE) in contextual bandits with finite action space. Inverse Propensity Score (IPS) weighting is a widely used method for OPE due to its unbiased…