247 citations · 379 across the 29 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2024★ 8 cited
Secrets of RLHF in Large Language Models Part II: Reward Modeling
Binghai Wang, Rui Zheng, Lu Chen +24
Reinforcement Learning from Human Feedback (RLHF) has become a crucial technology for aligning language models with human values and intentions, enabling models to produce more hel…
cs.AI2022
Towards Collaborative Question Answering: A Preliminary Study
Xiangkun Hu, Hang Yan, Qipeng Guo +3
Knowledge and expertise in the real-world can be disjointedly owned. To solve a complex question, collaboration among experts is often called for. In this paper, we propose CollabQ…