1 paper
Lirui Luo, Guoxi Zhang, Hongming Xu +3
Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes ca…