4 citations · 7 across the 3 of their papers we have counts for
6 papers
Design of Experiments for Stochastic Contextual Linear Bandits
Andrea Zanette, Kefan Dong, Jonathan Lee +1
In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, t…
Refined Analysis of FPL for Adversarial Markov Decision Processes
Yuanhao Wang, Kefan Dong
We consider the adversarial Markov Decision Process (MDP) problem, where the rewards for the MDP can be adversarially chosen, and the transition function can be either known or unk…
Multinomial Logit Bandit with Low Switching Cost
Kefan Dong, Yingkai Li, Qin Zhang +1
We study multinomial logit bandit with limited adaptivity, where the algorithms change their exploration actions as infrequently as possible when achieving almost optimal minimax r…
-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank
Kefan Dong, Jian Peng, Yining Wang +1
In this paper, we consider the problem of online learning of Markov decision processes (MDPs) with very large state spaces. Under the assumptions of realizable function approximati…
Exploration via Hindsight Goal Generation
Zhizhou Ren, Kefan Dong, Yuan Zhou +2
Goal-oriented reinforcement learning has recently been a practical framework for robotic manipulation tasks, in which an agent is required to reach a certain goal defined by a func…
Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP
Kefan Dong, Yuanhao Wang, Xiaoyu Chen +1
A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. \cite{jin2018q} proposed a Q-learning algorithm with UC…