collaborators

13 papers

cs.AI2026

Difficulty-Aware Semantic-ID Optimization for Generative Recommendation

Xin Yu, Stephen Li, Sina Aghaei +6

Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over hierarchical item identifiers. A common recipe is SFT followed by GRPO, ye…

stat.ML2026

Q-Learning with Fine-Grained Gap-Dependent Regret

Haochen Zhang, Zhong Zheng, Lingzhou Xue

We study fine-grained gap-dependent regret bounds for model-free reinforcement learning in episodic tabular Markov Decision Processes. Existing model-free algorithms achieve minima…

cs.LG2026

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Xin Yu, Liuchen Liao, Yiwen Zhang +3

On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its reliance on a stronger external teacher has…

stat.ML2026

Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement Learning

Haochen Zhang, Zhong Zheng, Lingzhou Xue

Motivated by real-world settings where data collection and policy deployment -- whether for a single agent or across multiple agents -- are costly, we study the problem of on-polic…

stat.ML2026

Gap-Dependent Bounds for Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation

Haochen Zhang, Zhong Zheng, Lingzhou Xue

We study gap-dependent performance guarantees for nearly minimax-optimal algorithms in reinforcement learning with linear function approximation. While prior works have established…

math.OC2026

Adaptive Algorithms for Robust Phase Retrieval

Zhong Zheng, Necdet Serhat Aybat, Shiqian Ma +1

This paper considers the robust phase retrieval, which can be cast as a nonsmooth and nonconvex composite optimization problem. We propose two first-order algorithms with adaptive…