1 paper
Guanqun Zhao, Zijun Xie, Binbin Zheng +3
Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior poli…