1 paper
Joe Suk, Jung-hun Kim
We study an infinite-armed bandit problem where actions' mean rewards are initially sampled from a reservoir distribution. Most prior works in this setting focused on stationary re…