1 paper
Yuchen He, Baolong Bi, Shenghua Liu +7
Although multi-objective reinforcement learning (MORL) is central to aligning large language models with complex human preferences, the prevailing practice of static weighted summa…