3 citations · 3 across the 1 of their papers we have counts for
2 papers
cs.AI2024★ 3 cited
Direct Language Model Alignment from Online AI Feedback
Shangmin Guo, Biao Zhang, Tianlin Liu +9
Direct alignment from preferences (DAP) methods, such as DPO, have recently emerged as efficient alternatives to reinforcement learning from human feedback (RLHF), that do not requ…
cs.LG2024
Decoding-time Realignment of Language Models
Tianlin Liu, Shangmin Guo, Leonardo Bianco +7
Aligning language models with human preferences is crucial for reducing errors and biases in these models. Alignment techniques, such as reinforcement learning from human feedback…