94 citations · 94 across the 5 of their papers we have counts for
1 paper · 1 filter
Miaomiao Ji, Yanqiu Wu, Zhibin Wu +4
Reward design plays a pivotal role in aligning large language models (LLMs) with human values, serving as the bridge between feedback signals and model optimization. This survey pr…