1 paper · 1 filter
Yufeng Zhang, Liyu Chen, Boyi Liu +4
Recent advances in reinforcement learning (RL) algorithms aim to enhance the performance of language models at scale. Yet, there is a noticeable absence of a cost-effective and sta…