4 papers
Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching
Preston Fu, Kevin Frans, Oleh Rybkin +2
Current paradigms for training language models via reinforcement learning rely heavily on sparse outcome rewards. However, as we pursue tasks that require longer and more complicat…
Compute-Optimal Scaling for Value-Based Deep RL
Preston Fu, Oleh Rybkin, Zhiyuan Zhou +4
As models grow larger and training them becomes expensive, it becomes increasingly important to scale training recipes not just to larger models and more data, but to do so in a co…
Value-Based Deep RL Scales Predictably
Oleh Rybkin, Michal Nauman, Preston Fu +4
Scaling data and compute is critical to the success of modern ML. However, scaling demands predictability: we want methods to not only perform well with more compute or data, but a…
Reasoning and Tools for Human-Level Forecasting
Elvis Hsieh, Preston Fu, Jonathan Chen
Language models (LMs) trained on web-scale datasets are largely successful due to their ability to memorize large amounts of training data, even if only present in a few examples.…