1 paper
Koichi Tanaka, Ren Kishimoto, Bushun Kawagishi +4
We study off-policy learning (OPL) in contextual bandits, which plays a key role in a wide range of real-world applications such as recommendation systems and online advertising. T…