Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
Chenghao Yang, Lin Gui, Chenxiao Yang +3
Reinforcement learning with verifiable rewards (RLVR) is a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs), yet its success hinges on eff…
cs.CL2024
BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
Lin Gui, Cristina Gârbacea, Victor Veitch
This paper concerns the problem of aligning samples from large language models to human preferences using best-of- sampling, where we draw samples, rank them, and return the…