1 paper
Yizhou Chen, Yawen Liu, Xuesi Wang +5
The reward model (RM) that represents human preferences plays a crucial role in optimizing the outputs of large language models (LLMs), e.g., through reinforcement learning from hu…