1 paper
Ryoma Sato, Shinji Ito
While classical formulations of multi-armed bandit problems assume that each arm's reward is independent and stationary, real-world applications often involve non-stationary enviro…