2 papers
cs.LG2026
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
Chanuk Lee, Minki Kang, Sung Ju Hwang
Recent studies observe that reinforcement learning with verifiable rewards (RLVR) reliably improves pass@1 on reasoning tasks, yet often fails to yield comparable gains in pass@k,…
cs.AI2026
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
Chanuk Lee, Sangwoo Park, Minki Kang +1
Reinforcement learning with verifiable rewards (RLVR) has emerged as a scalable paradigm for improving the reasoning capabilities of large language models. However, its effectivene…