1 paper · 1 filter
Hiroshi Takahashi, Tomoharu Iwata, Atsutoshi Kumagai +4
Aligning language models with human preferences is essential for ensuring their safety and reliability. Although most existing approaches assume specific human preference models su…