1 paper · 1 filter
Zhuo Li, Yuege Feng, Dandan Guo +3
The reward model (RM) plays a crucial role in aligning Large Language Models (LLMs) with human preferences through Reinforcement Learning, where the Bradley-Terry (BT) objective ha…