Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Self-Improving Robust Preference Optimization
Eugene Choi, Arash Ahmadian, Matthieu Geist +2
Online and offline RLHF methods, such as PPO and DPO, have been highly successful in aligning AI with human preferences. Despite their success, however, these methods suffer from f…
cs.LG2025
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub +8
Reinforcement Learning (RL) has been used to finetune Large Language Models (LLMs) using a reward model trained from preference data, to better align with human judgment. The recen…