8 citations · 13 across the 11 of their papers we have counts for
1 paper · 2 filters
Leitian Tao, Xuefeng Du, Sharon Li
Reward modeling, crucial for aligning large language models (LLMs) with human preferences, is often bottlenecked by the high cost of preference data. Existing textual data synthesi…