activity
20172022
most citedModel-Based Reinforcement Learning with Value-Targeted Regression

69 citations · 365 across the 21 of their papers we have counts for

collaborators

39 papers

cs.LG20221 cited

Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Jinghan Wang, Mengdi Wang, Lin F. Yang

This work considers the sample complexity of obtaining an -optimal policy in an average reward Markov Decision Process (AMDP), given access to a generative model (simu…

cs.LG20213 cited

Provably Breaking the Quadratic Error Compounding Barrier in Imitation Learning, Optimally

Nived Rajaraman, Yanjun Han, Lin F. Yang +2

We study the statistical limits of Imitation Learning (IL) in episodic Markov Decision Processes (MDPs) with a state space . We focus on the known-transition setting w…

cs.LG20217 cited

A Provably Efficient Algorithm for Linear Markov Decision Process with Low Switching Cost

Minbo Gao, Tianle Xie, Simon S. Du +1

Many real-world applications, such as those in medical domains, recommendation systems, etc, can be formulated as large state space reinforcement learning problems with only a smal…

cs.LG20205 cited

Minimax Sample Complexity for Turn-based Stochastic Game

Qiwen Cui, Lin F. Yang

The empirical success of Multi-agent reinforcement learning is encouraging, while few theoretical guarantees have been revealed. In this work, we prove that the plug-in solver appr…

cs.LG2020

Accommodating Picky Customers: Regret Bound and Exploration Complexity for Multi-Objective Reinforcement Learning

Jingfeng Wu, Vladimir Braverman, Lin F. Yang

In this paper we consider multi-objective reinforcement learning where the objectives are balanced using preferences. In practice, the preferences are often given in an adversarial…

cs.LG20204 cited

Episodic Linear Quadratic Regulators with Low-rank Transitions

Tianyu Wang, Lin F. Yang

Linear Quadratic Regulators (LQR) achieve enormous successful real-world applications. Very recently, people have been focusing on efficient learning algorithms for LQRs when their…