Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
Yiran Shen, Yu Xia, Jonathan Chang +1
Aligning large language models to human preferences is inherently multidimensional, yet most pipelines collapse heterogeneous signals into a single objective. We seek to answer wha…
cs.LG2024
Critique-out-Loud Reward Models
Zachary Ankner, Mansheej Paul, Brandon Cui +2
Traditionally, reward models used for reinforcement learning from human feedback (RLHF) are trained to directly predict preference scores without leveraging the generation capabili…