8 citations · 9 across the 2 of their papers we have counts for
7 papers
Near-Optimal Regret Bounds for Multi-batch Reinforcement Learning
Zihan Zhang, Yuhang Jiang, Yuan Zhou +1
In this paper, we study the episodic reinforcement learning (RL) problem modeled by finite-horizon Markov Decision Processes (MDPs) with constraint on the number of batches. The mu…
Improving Non-autoregressive Generation with Mixup Training
Ting Jiang, Shaohan Huang, Zihan Zhang +6
While pre-trained language models have achieved great success on various natural language understanding tasks, how to effectively leverage them into non-autoregressive generation t…
Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP
Zihan Zhang, Jiaqi Yang, Xiangyang Ji +1
This paper presents new \emph{variance-aware} confidence sets for linear bandits and linear mixture Markov Decision Processes (MDPs). With the new confidence sets, we obtain the fo…
Nearly Minimax Optimal Reward-free Reinforcement Learning
Zihan Zhang, Simon S. Du, Xiangyang Ji
We study the reward-free reinforcement learning framework, which is particularly suitable for batch reinforcement learning and scenarios where one needs policies for multiple rewar…
Model-Free Reinforcement Learning: from Clipped Pseudo-Regret to Sample Complexity
Zihan Zhang, Yuan Zhou, Xiangyang Ji
In this paper we consider the problem of learning an -optimal policy for a discounted Markov Decision Process (MDP). Given an MDP with states, actions, the discount fact…
Almost Optimal Model-Free Reinforcement Learning via Reference-Advantage Decomposition
Zihan Zhang, Yuan Zhou, Xiangyang Ji
We study the reinforcement learning problem in the setting of finite-horizon episodic Markov Decision Processes (MDPs) with states, actions, and episode length . We prop…