1 citations · 1 across the 1 of their papers we have counts for
1 paper
Tianci Xue, Ziqi Wang, Heng Ji
Aligning large language models (LLMs) with human preferences is essential for safe and useful LLMs. Previous works mainly adopt reinforcement learning (RLHF) and direct preference…