activity
20222026
most citedPrompt Learning for News Recommendation

77 citations · 118 across the 11 of their papers we have counts for

collaborators

11 papers

cs.AI2026

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Zhiqin Yang, Jingwen Fu, Yuhan Liu +16

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code,…

cs.LG2026

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

Xinmu Ge, Zizhuo Zhang, Yu Huang +9

On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. Under the reverse KL objective, the idealized optimum of OPD aligns the…

cs.LG2025

What Is Preference Optimization Doing, and Why?

Yue Wang, Qizhou Wang, Zizhuo Zhang +3

Preference optimization (PO) is indispensable for large language models (LLMs), with methods such as direct preference optimization (DPO) and proximal policy optimization (PPO) ach…

cs.LG2025

Towards Understanding Valuable Preference Data for Large Language Model Alignment

Zizhuo Zhang, Qizhou Wang, Shanshan Ye +4

Large language model (LLM) alignment is typically achieved through learning from human preference comparisons, making the quality of preference data critical to its success. Existi…

cs.LG2025

Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models

Zizhuo Zhang, Jianing Zhu, Xinmu Ge +6

While reinforcement learning with verifiable rewards (RLVR) is effective to improve the reasoning ability of large language models (LLMs), its reliance on human-annotated labels le…

cs.LG2025

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment

Jiazheng Zhang, Wenqing Jing, Zizhuo Zhang +9

Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human values. However, noisy preferences in human feedback can lead to reward misgeneralizatio…