13 citations · 14 across the 15 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration
Yan Sun, Jia Guo, Stanley Kok +3
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning ability of large language models, yet training remains costly because many rollouts contribute litt…
cs.LG2025
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Xinyu Tang, Zhenduo Zhang, Yurou Liu +4
Recent advances in large reasoning models have leveraged reinforcement learning with verifiable rewards (RLVR) to improve reasoning capabilities. However, scaling these methods typ…