1 paper · 1 filter
Yuanhong Wu, Djallel Bouneffouf, D. Frank Hsu
Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants ha…