1 paper
Youngmin Oh, Jinje Park, Taejin Paik
We introduce the first variance-aware algorithms for contextual dueling bandits that leverage shallow exploration strategies with neural networks for nonlinear utility approximatio…