1 paper
Renjie Mao, Xiangxin Zhou, Lvfang Tao +7
Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing PPO-style trust-region mechanisms remain position-agnostic…