1 paper
Haw-Shiuan Chang, Jeffrey Gomez, Mehul Patwari +2
To align a Large Language Model (LLM), most existing methods collect explicit human feedback and train a reward model to predict the human preference based on the response text. Th…