2 citations · 2 across the 5 of their papers we have counts for
1 paper · 1 filter
Xingjian Wu, Junlin Liu, Xingchen Liu +6
Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). Whil…