1 paper
Timo Kaufmann, Yannick Metz, Daniel Keim +1
Binary choices, as often used for reinforcement learning from human feedback (RLHF), convey only the direction of a preference. A person may choose apples over oranges and bananas…