activity
20202026
most citedJellyfish: A Large Language Model for Data Preprocessing

4 citations · 7 across the 25 of their papers we have counts for

collaborators
Showing stat.MLShow all

7 papers · 1 filter

stat.ML2026

Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost

Zihang Liang, Haochen Zhang, Lingzhou Xue

We study federated online reinforcement learning with linear function approximation. While recent multi-agent reinforcement learning algorithms achieve strong regret guarantees, th…

stat.ML2026

Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation

Haochen Zhang, Zhong Zheng, Lingzhou Xue

We study gap-dependent performance guarantees for nearly minimax-optimal algorithms in reinforcement learning with linear function approximation. While prior works have established…

stat.ML2025

Q-Learning with Fine-Grained Gap-Dependent Regret

Haochen Zhang, Zhong Zheng, Lingzhou Xue

We study fine-grained gap-dependent regret bounds for model-free reinforcement learning in episodic tabular Markov Decision Processes. Existing model-free algorithms achieve minima…

stat.ML2025

Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning

Haochen Zhang, Zhong Zheng, Lingzhou Xue

Motivated by real-world settings where data collection and policy deployment -- whether for a single agent or across multiple agents -- are costly, we study the problem of on-polic…

stat.ML2025

Gap-Dependent Bounds for Federated -learning

Haochen Zhang, Zhong Zheng, Lingzhou Xue

We present the first gap-dependent analysis of regret and communication cost for on-policy federated -Learning in tabular episodic finite-horizon Markov decision processes (MDPs…

stat.ML2024

Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition

Zhong Zheng, Haochen Zhang, Lingzhou Xue

We study the gap-dependent bounds of two important algorithms for on-policy Q-learning for finite-horizon episodic tabular Markov Decision Processes (MDPs): UCB-Advantage (Zhang et…