23 citations · 24 across the 2 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024★ 1 cited
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
Karel D'Oosterlinck, Winnie Xu, Chris Develder +5
Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes…
cs.LG2024★ 23 cited
KTO: Model Alignment as Prospect Theoretic Optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff +2
Kahneman & Tversky's tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-ave…