1 paper
Xiuyu Li, Jinkai Zhang, Mingyang Yi +4
Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complicated training process. To addres…