1 paper
Jing Zhao, Ting Zhen, Junwei Bao +2
Current alignment methods for Large Language Models (LLMs) rely on compressing vast amounts of human preference data into static, absolute reward functions, leading to data scarcit…