activity
20192022
most citedImproving Non-autoregressive Generation with Mixup Training

8 citations · 9 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG20221 cited

Near-Optimal Regret Bounds for Multi-batch Reinforcement Learning

Zihan Zhang, Yuhang Jiang, Yuan Zhou +1

In this paper, we study the episodic reinforcement learning (RL) problem modeled by finite-horizon Markov Decision Processes (MDPs) with constraint on the number of batches. The mu…

cs.CL20218 cited

Improving Non-autoregressive Generation with Mixup Training

Ting Jiang, Shaohan Huang, Zihan Zhang +6

While pre-trained language models have achieved great success on various natural language understanding tasks, how to effectively leverage them into non-autoregressive generation t…

cs.LG2021

Improved Variance-Aware Confidence Sets for Linear Bandits and Linear Mixture MDP

Zihan Zhang, Jiaqi Yang, Xiangyang Ji +1

This paper presents new \emph{variance-aware} confidence sets for linear bandits and linear mixture Markov Decision Processes (MDPs). With the new confidence sets, we obtain the fo…

cs.LG2020

Nearly Minimax Optimal Reward-free Reinforcement Learning

Zihan Zhang, Simon S. Du, Xiangyang Ji

We study the reward-free reinforcement learning framework, which is particularly suitable for batch reinforcement learning and scenarios where one needs policies for multiple rewar…

cs.LG2020

Model-Free Reinforcement Learning: from Clipped Pseudo-Regret to Sample Complexity

Zihan Zhang, Yuan Zhou, Xiangyang Ji

In this paper we consider the problem of learning an -optimal policy for a discounted Markov Decision Process (MDP). Given an MDP with states, actions, the discount fact…

cs.LG2020

Almost Optimal Model-Free Reinforcement Learning via Reference-Advantage Decomposition

Zihan Zhang, Yuan Zhou, Xiangyang Ji

We study the reinforcement learning problem in the setting of finite-horizon episodic Markov Decision Processes (MDPs) with states, actions, and episode length . We prop…