4 citations · 8 across the 28 of their papers we have counts for
1 paper · 1 filter
Ganqu Cui, Lifan Yuan, Ning Ding +9
Learning from human feedback has become a pivot technique in aligning large language models (LLMs) with human preferences. However, acquiring vast and premium human feedback is bot…