1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.LG2025★ 1 cited
Outcome-based Exploration for LLM Reasoning
Yuda Song, Julia Kempe, Remi Munos
Reinforcement learning (RL) has emerged as a powerful method for improving the reasoning abilities of large language models (LLMs). Outcome-based RL, which rewards policies solely…
cs.LG2025
Accelerating Unbiased LLM Evaluation via Synthetic Feedback
Zhaoyi Zhou, Yuda Song, Andrea Zanette
When developing new large language models (LLMs), a key step is evaluating their final performance, often by computing the win-rate against a reference model based on external feed…