2 papers
cs.AI2024
Towards Comprehensive Preference Data Collection for Reward Modeling
Yulan Hu, Qingyang Li, Sheng Ouyang +6
Reinforcement Learning from Human Feedback (RLHF) facilitates the alignment of large language models (LLMs) with human preferences, thereby enhancing the quality of responses gener…
cs.CL2023
KwaiYiiMath: Technical Report
Jiayi Fu, Lei Lin, Xiaoyang Gao +18
Recent advancements in large language models (LLMs) have demonstrated remarkable abilities in handling a variety of natural language processing (NLP) downstream tasks, even on math…