activity
20242026
collaborators
Showing cs.LGShow all

15 papers · 1 filter

cs.LG2026

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

Runlong Zhou, Zihan Zhang, Maryam Fazel +1

We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with states, actions, horizon , and per-trajectory total…

cs.LG2025

Convergence Dynamics of Over-Parameterized Score Matching for a Single Gaussian

Yiran Zhang, Weihang Xu, Mo Zhou +2

Score matching has become a central training objective in modern generative modeling, particularly in diffusion models, where it is used to learn high-dimensional data distribution…

cs.LG2025

Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback

Runlong Zhou, Maryam Fazel, Simon S. Du

Reinforcement learning from human feedback (RLHF) has become essential for improving language model capabilities, but traditional approaches rely on the assumption that human prefe…

cs.LG2025

Policy-Based Trajectory Clustering in Offline Reinforcement Learning

Hao Hu, Xinqi Wang, Simon Shaolei Du

We introduce a novel task of clustering trajectories from offline reinforcement learning (RL) datasets, where each cluster center represents the policy that generated its trajector…

cs.LG2025

Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures

Mo Zhou, Weihang Xu, Maryam Fazel +1

Learning Gaussian Mixture Models (GMMs) is a fundamental problem in statistics and machine learning, with the Expectation-Maximization (EM) algorithm and its popular variant gradie…

cs.LG2025

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

Shulun Chen, Runlong Zhou, Zihan Zhang +2

We consider the gap-dependent regret bounds for episodic MDPs. We show that the Monotonic Value Propagation (MVP) algorithm achieves a variance-aware gap-dependent regret bound of…