7 citations · 18 across the 6 of their papers we have counts for
4 papers · 1 filter
How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR
Richa Verma, Balaraman Ravindran
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for reasoning in language models, with GRPO as its primary example. However, GRPO requires…
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
Richa Verma, Bavish Kulur, Sanjay Chawla +1
We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it from scratch. While costs cou…
SIBRE: Self Improvement Based REwards for Adaptive Feedback in Reinforcement Learning
Somjit Nath, Richa Verma, Abhik Ray +1
We propose a generic reward shaping approach for improving the rate of convergence in reinforcement learning (RL), called Self Improvement Based REwards, or SIBRE. The approach is…
Accelerating Training in Pommerman with Imitation and Reinforcement Learning
Hardik Meisheri, Omkar Shelke, Richa Verma +1
The Pommerman simulation was recently developed to mimic the classic Japanese game Bomberman, and focuses on competitive gameplay in a multi-agent setting. We focus on the 2$\times…