2 citations · 4 across the 17 of their papers we have counts for
4 papers · 1 filter
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
Congmin Zheng, Jiachen Zhu, Jianghao Lin +6
Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. Howeve…
Process Reward Model with Q-Value Rankings
Wendi Li, Yixuan Li
Process Reward Modeling (PRM) is critical for complex reasoning and decision-making tasks where the accuracy of intermediate steps significantly influences the overall outcome. Exi…
Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue
Shixuan Fan, Wei Wei, Wendi Li +3
The core of the dialogue system is to generate relevant, informative, and human-like responses based on extensive dialogue history. Recently, dialogue generation domain has seen ma…
Reinforcement Learning with Token-level Feedback for Controllable Text Generation
Wendi Li, Wei Wei, Kaihe Xu +3
To meet the requirements of real-world applications, it is essential to control generations of large language models (LLMs). Prior research has tried to introduce reinforcement lea…