activity
20182022
most citedOnline Model Selection for Reinforcement Learning with Function Approximation

3 citations · 9 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG20221 cited

Oracle Inequalities for Model Selection in Offline Reinforcement Learning

Jonathan N. Lee, George Tucker, Ofir Nachum +2

In offline reinforcement learning (RL), a learner leverages prior logged data to learn a good policy without interacting with the environment. A major challenge in applying such me…

cs.LG20213 cited

Design of Experiments for Stochastic Contextual Linear Bandits

Andrea Zanette, Kefan Dong, Jonathan Lee +1

In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, t…

cs.LG20211 cited

Near Optimal Policy Optimization via REPS

Aldo Pacchiano, Jonathan Lee, Peter Bartlett +1

Since its introduction a decade ago, \emph{relative entropy policy search} (REPS) has demonstrated successful policy learning on a number of simulated and real-world robotic domain…

cs.LG20203 cited

Online Model Selection for Reinforcement Learning with Function Approximation

Jonathan N. Lee, Aldo Pacchiano, Vidya Muthukumar +2

Deep reinforcement learning has achieved impressive successes yet often requires a very large amount of interaction data. This result is perhaps unsurprising, as using complicated…

cs.LG20201 cited

Is Q-Learning Provably Efficient? An Extended Analysis

Kushagra Rastogi, Jonathan Lee, Fabrice Harel-Canada +1

This work extends the analysis of the theoretical results presented within the paper Is Q-Learning Provably Efficient? by Jin et al. We include a survey of related research to cont…

cs.LG2020

Accelerated Message Passing for Entropy-Regularized MAP Inference

Jonathan N. Lee, Aldo Pacchiano, Peter Bartlett +1

Maximum a posteriori (MAP) inference in discrete-valued Markov random fields is a fundamental problem in machine learning that involves identifying the most likely configuration of…