2 papers
cs.LG2025
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Jingcheng Hu, Yinmin Zhang, Qi Han +3
We introduce Open-Reasoner-Zero, the first open source implementation of large-scale reasoning-oriented RL training on the base model focusing on scalability, simplicity and access…
cs.CL2024
Xwin-LM: Strong and Scalable Alignment Practice for LLMs
Bolin Ni, JingCheng Hu, Yixuan Wei +4
In this work, we present Xwin-LM, a comprehensive suite of alignment methodologies for large language models (LLMs). This suite encompasses several key techniques, including superv…