4 citations · 8 across the 20 of their papers we have counts for
1 paper · 2 filters
Zhaowei Zhang, Xiaohan Liu, Xuekai Zhu +6
Reinforcement learning with verifiable rewards (RLVR) has achieved remarkable success in logical reasoning tasks, yet whether large language model (LLM) alignment requires fundamen…