1 paper · 1 filter
Khaoula Chehbouni, Jonathan Colaço Carr, Yash More +2
In an effort to mitigate the harms of large language models (LLMs), learning from human feedback (LHF) has been used to steer LLMs towards outputs that are intended to be both less…