77 citations · 118 across the 11 of their papers we have counts for
11 papers
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
Zhiqin Yang, Jingwen Fu, Yuhan Liu +16
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code,…
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
Xinmu Ge, Zizhuo Zhang, Yu Huang +9
On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. Under the reverse KL objective, the idealized optimum of OPD aligns the…
What Is Preference Optimization Doing, and Why?
Yue Wang, Qizhou Wang, Zizhuo Zhang +3
Preference optimization (PO) is indispensable for large language models (LLMs), with methods such as direct preference optimization (DPO) and proximal policy optimization (PPO) ach…
Towards Understanding Valuable Preference Data for Large Language Model Alignment
Zizhuo Zhang, Qizhou Wang, Shanshan Ye +4
Large language model (LLM) alignment is typically achieved through learning from human preference comparisons, making the quality of preference data critical to its success. Existi…
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
Zizhuo Zhang, Jianing Zhu, Xinmu Ge +6
While reinforcement learning with verifiable rewards (RLVR) is effective to improve the reasoning ability of large language models (LLMs), its reliance on human-annotated labels le…
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
Jiazheng Zhang, Wenqing Jing, Zizhuo Zhang +9
Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human values. However, noisy preferences in human feedback can lead to reward misgeneralizatio…