1 paper
Zeyu Jia, Jian Qian, Alexander Rakhlin +1
We consider realizable contextual bandits with general function approximation, investigating how small reward variance can lead to better-than-minimax regret bounds. Unlike in mini…