1 paper
Jingcheng Hu, Yinmin Zhang, Qi Han +3
We introduce Open-Reasoner-Zero, the first open source implementation of large-scale reasoning-oriented RL training on the base model focusing on scalability, simplicity and access…