3 citations · 9 across the 6 of their papers we have counts for
Showing 2021Show all
2 papers · 1 filter
cs.LG2021★ 3 cited
Design of Experiments for Stochastic Contextual Linear Bandits
Andrea Zanette, Kefan Dong, Jonathan Lee +1
In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, t…
cs.LG2021★ 1 cited
Near Optimal Policy Optimization via REPS
Aldo Pacchiano, Jonathan Lee, Peter Bartlett +1
Since its introduction a decade ago, \emph{relative entropy policy search} (REPS) has demonstrated successful policy learning on a number of simulated and real-world robotic domain…