4 papers
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
Fan Chen, Zeyu Jia, Alexander Rakhlin +1
Reinforcement learning with outcome-based feedback faces a fundamental challenge: when rewards are only observed at trajectory endpoints, how do we assign credit to the right actio…
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
Yurun Yuan, Fan Chen, Zeyu Jia +2
Policy-based methods currently dominate reinforcement learning (RL) pipelines for large language model (LLM) reasoning, leaving value-based approaches largely unexplored. We revisi…
Beyond Covariance Matrix: The Statistical Complexity of Private Linear Regression
Fan Chen, Jiachun Li, Alexander Rakhlin +1
We study the statistical complexity of private linear regression under an unknown, potentially ill-conditioned covariate distribution. Somewhat surprisingly, under privacy constrai…
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
Fan Chen, Alexander Rakhlin
We study the problem of interactive decision making in which the underlying environment changes over time subject to given constraints. We propose a framework, which we call \texti…