6 papers
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
Yurun Yuan, Fan Chen, Zeyu Jia +2
Policy-based methods currently dominate reinforcement learning (RL) pipelines for large language model (LLM) reasoning, leaving value-based approaches largely unexplored. We revisi…
A Gapped Scale-Sensitive Dimension and Lower Bounds for Offset Rademacher Complexity
Zeyu Jia, Yury Polyanskiy, Alexander Rakhlin
We study gapped scale-sensitive dimensions of a function class in both sequential and non-sequential settings. We demonstrate that covering numbers for any uniformly bounded class…
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
Fan Chen, Zeyu Jia, Alexander Rakhlin +1
Reinforcement learning with outcome-based feedback faces a fundamental challenge: when rewards are only observed at trajectory endpoints, how do we assign credit to the right actio…
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
Zeyu Jia, Alexander Rakhlin, Tengyang Xie
As large language models have evolved, it has become crucial to distinguish between process supervision and outcome supervision -- two key reinforcement learning approaches to comp…
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
Zeyu Jia, Yury Polyanskiy, Alexander Rakhlin
We study the problem of sequential probability assignment under logarithmic loss, both with and without side information. Our objective is to analyze the minimax regret -- a notion…
How Does Variance Shape the Regret in Contextual Bandits?
Zeyu Jia, Jian Qian, Alexander Rakhlin +1
We consider realizable contextual bandits with general function approximation, investigating how small reward variance can lead to better-than-minimax regret bounds. Unlike in mini…