activity
20242026
collaborators

17 papers

cs.LG2026

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

Runlong Zhou, Zihan Zhang, Maryam Fazel +1

We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with states, actions, horizon , and per-trajectory total…

cs.CL2026

Cold-Start Personalization via Training-Free Priors from Structured World Models

Avinandan Bose, Shuyue Stella Li, Faeze Brahman +6

Cold-start personalization requires inferring user preferences through interaction when no user-specific historical data is available. The core challenge is a routing problem: each…

cs.LG2025

Convergence Dynamics of Over-Parameterized Score Matching for a Single Gaussian

Yiran Zhang, Weihang Xu, Mo Zhou +2

Score matching has become a central training objective in modern generative modeling, particularly in diffusion models, where it is used to learn high-dimensional data distribution…

math.OC2025

Global Convergence of Four-Layer Matrix Factorization under Random Initialization

Minrui Luo, Weihang Xu, Xiang Gao +2

Gradient descent dynamics on the deep matrix factorization problem is extensively studied as a simplified theoretical model for deep neural networks. Although the convergence theor…

cs.LG2025

Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback

Runlong Zhou, Maryam Fazel, Simon S. Du

Reinforcement learning from human feedback (RLHF) has become essential for improving language model capabilities, but traditional approaches rely on the assumption that human prefe…

cs.LG2025

Policy-Based Trajectory Clustering in Offline Reinforcement Learning

Hao Hu, Xinqi Wang, Simon Shaolei Du

We introduce a novel task of clustering trajectories from offline reinforcement learning (RL) datasets, where each cluster center represents the policy that generated its trajector…