3 citations · 7 across the 10 of their papers we have counts for
1 paper · 2 filters
Miao Peng, Weizhou Shen, Nuo Chen +3
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective in enhancing LLMs short-context reasoning, but its performance degrades in long-context scenarios that re…