budget allocation 1optimal stopping 1reinforcement learning 1sample efficiency 1sequential decision making 1
From the 1 of 2 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
PAIR: Pairwise-Aware Inclusion Reweighting for Adaptive Rollout Allocation in RLVR
Pixel Nomand, Elena Voss, Marcus Hale +1
Reinforcement learning with verifiable rewards (RLVR) spends most of its compute generating groups of long reasoning trajectories. Recent allocators reduce this cost by assigning b…
cs.LG2026
Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR
Pixel Nomand, Elena Voss, Marcus Hale +1
The paper proposes SARA, a method that adaptively stops rollouts for reinforcement learning with verifiable rewards by using a Bayesian model and sequential hypothesis testing to s…