1 paper · 1 filter
Yuanhong Wu, Djallel Bouneffouf, D. Frank Hsu
Aligning large language models (LLMs) with human values is a central challenge for ensuring trustworthy and safe deployment. While existing methods such as Reinforcement Learning f…