Showing 2026Show all
3 papers · 1 filter
cs.LG2026
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
Stanley Wei, Juno Kim
Recent advances in large language models (LLMs) have demonstrated that reinforcement fine-tuning of pretrained base models can lead to significant gains in reasoning performance at…
cs.LG2026
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
Shuning Shang, Hubert Strauss, Stanley Wei +2
Training language models via reinforcement learning often relies on imperfect proxy rewards, since ground truth rewards that precisely define the intended behavior are rarely avail…
cs.LG2026
Improved high-dimensional estimation with Langevin dynamics and stochastic weight averaging
Stanley Wei, Alex Damian, Jason D. Lee
Significant recent work has studied the ability of gradient descent to recover a hidden planted direction in different high-dimensional settings, including te…