1 citations · 3 across the 9 of their papers we have counts for
1 paper · 1 filter
Chen Jia
Preference learning (PL) with large language models (LLMs) aims to align the LLMs' generations with human preferences. Previous work on reinforcement learning from human feedback (…