1 paper
Zhendong Shi, Ercan E. Kuruoglu, Xiaoli Wei
In algorithm optimization in reinforcement learning, how to deal with the exploration-exploitation dilemma is particularly important. Multi-armed bandit problem can optimize the pr…