1 paper · 1 filter
Junxi Yin, Haisen Luo, Zhenyu Li +4
While Reinforcement Learning with Verifiable Rewards (RLVR) enhances complex reasoning in LLMs, current methods struggle to balance exploration and exploitation. This leads to crit…