2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Haotian Liu, Yihao Liu, Jingwei Ni +11
As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultivate comprehensive capabilities. This shift demands new algorith…