1 paper
Ziang Song, Tianle Cai, Jason D. Lee +1
The extraordinary capabilities of large language models (LLMs) such as ChatGPT and GPT-4 are in part unleashed by aligning them with reward models that are trained on human prefere…