238 citations · 392 across the 8 of their papers we have counts for
4 papers · 1 filter
BOND: Aligning LLMs with Best-of-N Distillation
Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot +17
Reinforcement learning from human feedback (RLHF) is a key driver of quality and safety in state-of-the-art large language models. Yet, a surprisingly simple and strong inference-t…
WARP: On the Benefits of Weight Averaged Rewarded Policies
Alexandre Ramé, Johan Ferret, Nino Vieillard +7
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) by encouraging their generations to have high rewards, using a reward model trained on human p…
WARM: On the Benefits of Weight Averaged Reward Models
Alexandre Ramé, Nino Vieillard, Léonard Hussenot +4
Aligning large language models (LLMs) with human preferences through reinforcement learning (RLHF) can lead to reward hacking, where LLMs exploit failures in the reward model (RM)…
Offline Reinforcement Learning with On-Policy Q-Function Regularization
Laixi Shi, Robert Dadashi, Yuejie Chi +2
The core challenge of offline reinforcement learning (RL) is dealing with the (potentially catastrophic) extrapolation error induced by the distribution shift between the history d…