activity
20192021
most citedMultinomial Logit Bandit with Low Switching Cost

4 citations · 7 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG20213 cited

Design of Experiments for Stochastic Contextual Linear Bandits

Andrea Zanette, Kefan Dong, Jonathan Lee +1

In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, t…

cs.LG2020

Refined Analysis of FPL for Adversarial Markov Decision Processes

Yuanhao Wang, Kefan Dong

We consider the adversarial Markov Decision Process (MDP) problem, where the rewards for the MDP can be adversarially chosen, and the transition function can be either known or unk…

cs.LG20204 cited

Multinomial Logit Bandit with Low Switching Cost

Kefan Dong, Yingkai Li, Qin Zhang +1

We study multinomial logit bandit with limited adaptivity, where the algorithms change their exploration actions as infrequently as possible when achieving almost optimal minimax r…

cs.LG2019

-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank

Kefan Dong, Jian Peng, Yining Wang +1

In this paper, we consider the problem of online learning of Markov decision processes (MDPs) with very large state spaces. Under the assumptions of realizable function approximati…

cs.LG2019

Exploration via Hindsight Goal Generation

Zhizhou Ren, Kefan Dong, Yuan Zhou +2

Goal-oriented reinforcement learning has recently been a practical framework for robotic manipulation tasks, in which an agent is required to reach a certain goal defined by a func…

cs.LG2019

Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP

Kefan Dong, Yuanhao Wang, Xiaoyu Chen +1

A fundamental question in reinforcement learning is whether model-free algorithms are sample efficient. Recently, Jin et al. \cite{jin2018q} proposed a Q-learning algorithm with UC…