2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Xinnan Zhang, Chenliang Li, Siliang Zeng +6
Aligning large language models (LLMs) with human preferences usually requires fine-tuning methods such as RLHF and DPO. These methods directly optimize the model parameters, so the…