2 papers
cs.LG2026
How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR
Richa Verma, Balaraman Ravindran
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for reasoning in language models, with GRPO as its primary example. However, GRPO requires…
cs.LG2026
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
Richa Verma, Bavish Kulur, Sanjay Chawla +1
We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it from scratch. While costs cou…