3 papers
cs.LG2026
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Yucheng Li, Huiqiang Jiang, Yang Xu +14
Reinforcement learning (RL) has become a key component in modern large language models, yet the rollout stage remains the key bottleneck in RL training pipelines. Although Multi-To…
cs.LG2025
Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking
Dhruv Rohatgi, Abhishek Shetty, Donya Saless +4
Test-time algorithms that combine the generative power of language models with process verifiers that assess the quality of partial generations offer a promising lever for elicitin…
cs.CL2025
On the Query Complexity of Verifier-Assisted Language Generation
Edoardo Botta, Yuchen Li, Aashay Mehta +3
Recently, a plethora of works have proposed inference-time algorithms (e.g. best-of-n), which incorporate verifiers to assist the generation process. Their quality-efficiency trade…