1 citations · 1 across the 7 of their papers we have counts for
4 papers · 1 filter
Verifier-Induced Support Reshaping in On-Policy Optimization
Shaohang Wei, Zikun Su, Feifan Song +4
We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sa…
Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu +7
Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing RLVR method…
Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning
Wenhao Yu, Shaohang Wei, Jiahong Liu +5
Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dimensional: the ground-truth probability…
One-Way Policy Optimization for Self-Evolving LLMs
Shuo Yang, Jinda Lu, Kexin Huang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become a promising paradigm for scaling reasoning capabilities of Large Language Models (LLMs). However, the sparsity of b…