1 paper
Julia Sepúlveda Coelho, Scott A. Hale
Large Language Models (LLMs) are often fine-tuned through Reinforcement Learning from Human Feedback (RLHF) to align with people's preferences and values. However, this method has…