2 citations · 2 across the 1 of their papers we have counts for
1 paper
Zhilin Wang, Yi Dong, Olivier Delalleau +6
High-quality preference datasets are essential for training reward models that can effectively guide large language models (LLMs) in generating high-quality responses aligned with…