1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Alexander Shypula, Shuo Li, Botong Zhang +3
Recent work suggests that preference-tuning techniques -- such as Reinforcement Learning from Human Feedback (RLHF) methods like PPO and GRPO, as well as alternatives like DPO -- r…