collaborators

6 papers

cs.LG2025

Process Reinforcement through Implicit Rewards

Ganqu Cui, Lifan Yuan, Zefan Wang +22

Dense process rewards have proven a more effective alternative to the sparse outcome-level rewards in the inference-time scaling of large language models (LLMs), particularly in ta…

cs.LG2025

RLPR: Extrapolating RLVR to General Domains without Verifiers

Tianyu Yu, Bo Ji, Shouli Wang +9

Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates promising potential in advancing the reasoning capabilities of LLMs. However, its success remains largely confine…

cs.CL2025

The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning

Bingxiang He, Ning Ding, Cheng Qian +10

Understanding alignment techniques begins with comprehending zero-shot generalization brought by instruction tuning, but little of the mechanism has been understood. Existing work…

cs.LG2024

Free Process Rewards without Process Labels

Lifan Yuan, Wendi Li, Huayu Chen +6

Different from its counterpart outcome reward models (ORMs), which evaluate the entire responses, a process reward model (PRM) scores a reasoning trajectory step by step, providing…

cs.LG2024

Noise Contrastive Alignment of Language Models with Explicit Rewards

Huayu Chen, Guande He, Lifan Yuan +3

User intentions are typically formalized as evaluation rewards to be maximized when fine-tuning language models (LMs). Existing alignment methods, such as Direct Preference Optimiz…

cs.CL2024

Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Yiju Guo, Ganqu Cui, Lifan Yuan +9

Alignment in artificial intelligence pursues the consistency between model responses and human preferences as well as values. In practice, the multifaceted nature of human preferen…