Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
The Bidirectional Process Reward Model
Lingyin Zhang, Jun Gao, Xiaoxue Ren +1
Process Reward Models (PRMs), which assign fine-grained scores to intermediate reasoning steps within a solution trajectory, have emerged as a promising approach to enhance the rea…
cs.CL2024
Guiding ChatGPT to Generate Salient Domain Summaries
Jun Gao, Ziqiang Cao, Shaoyao Huang +2
ChatGPT is instruct-tuned to generate general and human-expected content to align with human preference through Reinforcement Learning from Human Feedback (RLHF), meanwhile resulti…