1 citations · 1 across the 7 of their papers we have counts for
1 paper · 1 filter
Fei Wu, Zhenrong Zhang, Qikai Chang +3
Reinforcement Learning with Verifiable Rewards (RLVR) elicits long chain-of-thought reasoning in large language models (LLMs), but outcome-based rewards lead to coarse-grained adva…