Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR
Yash Ingle, Jaival Chauhan, Ankit Yadav +1
Reinforcement learning with verifiable rewards (RLVR) has become a highly effective method for improving the reasoning abilities of Large Language Models (LLMs). Recent research sh…
cs.LG2026
Selector-Guided Autonomous Curriculum for One-Shot Reinforcement Learning from Verifiable Rewards
Rudray Dave, Vedang Dubey, Smit Deoghare +1
Recently, Reinforcement Learning from Verifiable Rewards (RLVR) has been established as a highly effective technique for augmenting the math reasoning skills of Large Language Mode…