activity
20242026
most citedOutcome-based Exploration for LLM Reasoning

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Tail-Likelihood Reinforcement Learning

Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar +11

Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward whi…

cs.LG2026

Expanding the Capabilities of Reinforcement Learning via Text Feedback

Yuda Song, Lili Chen, Fahim Tajwar +5

The success of RL for LLM post-training stems from an unreasonably uninformative source: a single bit of information per rollout as binary reward or preference label. At the other…

cs.LG2026

Maximum Likelihood Reinforcement Learning

Fahim Tajwar, Guanning Zeng, Yueer Zhou +7

Reinforcement learning (RL) is the method of choice for training models in setups where the objective function can only be evaluated by sampling from the model. Our key observation…

cs.LG2025★ 1 cited

Outcome-based Exploration for LLM Reasoning

Yuda Song, Julia Kempe, Remi Munos

Reinforcement learning (RL) has emerged as a powerful method for improving the reasoning abilities of large language models (LLMs). Outcome-based RL, which rewards policies solely…

cs.LG2025

Accelerating Unbiased LLM Evaluation via Synthetic Feedback

Zhaoyi Zhou, Yuda Song, Andrea Zanette

When developing new large language models (LLMs), a key step is evaluating their final performance, often by computing the win-rate against a reference model based on external feed…

cs.LG2024

Rich-Observation Reinforcement Learning with Continuous Latent Dynamics

Yuda Song, Lili Wu, Dylan J. Foster +1

Sample-efficiency and reliability remain major bottlenecks toward wide adoption of reinforcement learning algorithms in continuous settings with high-dimensional perceptual inputs.…