activity
20242026
collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

High-accuracy sampling for diffusion models and log-concave distributions

Fan Chen, Sinho Chewi, Constantinos Daskalakis +1

We present algorithms for diffusion model sampling which obtain -error in steps, given access to -accurate score estimates in . Thi…

cs.LG2025

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits

Fan Chen, Zeyu Jia, Alexander Rakhlin +1

Reinforcement learning with outcome-based feedback faces a fundamental challenge: when rewards are only observed at trajectory endpoints, how do we assign credit to the right actio…

cs.LG2025

Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning

Yurun Yuan, Fan Chen, Zeyu Jia +2

Policy-based methods currently dominate reinforcement learning (RL) pipelines for large language model (LLM) reasoning, leaving value-based approaches largely unexplored. We revisi…

cs.LG2025

Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective

Zeyu Jia, Alexander Rakhlin, Tengyang Xie

As large language models have evolved, it has become crucial to distinguish between process supervision and outcome supervision -- two key reinforcement learning approaches to comp…

cs.LG2025

On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy

Zeyu Jia, Yury Polyanskiy, Alexander Rakhlin

We study the problem of sequential probability assignment under logarithmic loss, both with and without side information. Our objective is to analyze the minimax regret -- a notion…

cs.LG2025

Beyond Covariance Matrix: The Statistical Complexity of Private Linear Regression

Fan Chen, Jiachun Li, Alexander Rakhlin +1

We study the statistical complexity of private linear regression under an unknown, potentially ill-conditioned covariate distribution. Somewhat surprisingly, under privacy constrai…