4 citations · 4 across the 1 of their papers we have counts for
1 paper
Bram Wallace, Meihua Dang, Rafael Rafailov +7
Large language models (LLMs) are fine-tuned using human comparison data with Reinforcement Learning from Human Feedback (RLHF) methods to make them better aligned with users' prefe…