1 paper
Tan Zhu, Guannan Liang, Chunjiang Zhu +2
In stochastic contextual bandit (SCB) problems, an agent selects an action based on certain observed context to maximize the cumulative reward over iterations. Recently there have…