1 paper
Li Yang, Xiaodong Yan, Dandan Jiang
Multi-armed bandit (MAB) processes constitute a foundational subclass of reinforcement learning problems and represent a central topic in statistical decision theory. Yet, conducti…