3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CL2025
Discriminative Policy Optimization for Token-Level Reward Models
Hongzhan Chen, Tao Yang, Shiping Gao +4
Process reward models (PRMs) provide more nuanced supervision compared to outcome reward models (ORMs) for optimizing policy models, positioning them as a promising approach to enh…
cs.CL2024
FuseChat: Knowledge Fusion of Chat Models
Fanqi Wan, Longguang Zhong, Ziyi Yang +2
While training large language models (LLMs) from scratch can indeed lead to models with distinct capabilities and strengths, it incurs substantial costs and may lead to redundancy…
cs.CL2023★ 3 cited
Learning to Memorize Entailment and Discourse Relations for Persona-Consistent Dialogues
Ruijun Chen, Jin Wang, Liang-Chih Yu +1
Maintaining engagement and consistency is particularly important in dialogue systems. Existing works have improved the performance of dialogue systems by intentionally learning int…