activity
20232026
most citedLearning Optimal Advantage from Preferences and Mistaking it for Reward

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Mitigating Cognitive Bias in RLHF by Altering Rationality

Tiffany Horter, Andrew Markham, Niki Trigoni +1

How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outputs are used to train a rewar…

cs.LG2025

Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners

Calarina Muslimani, Kerrick Johnstonbaugh, Suyog Chandramouli +3

Reinforcement learning agents are fundamentally limited by the quality of the reward functions they learn from, yet reward design is often overlooked under the assumption that a we…

cs.LG2025

Influencing Humans to Conform to Preference Models for RLHF

Stephane Hatgis-Kessell, W. Bradley Knox, Serena Booth +1

Designing a reinforcement learning from human feedback (RLHF) algorithm to approximate a human's unobservable reward function requires assuming, implicitly or explicitly, a model o…

cs.CY2023

Quality-Diversity Generative Sampling for Learning with Synthetic Data

Allen Chang, Matthew C. Fontaine, Serena Booth +2

Generative models can serve as surrogates for some real data sources by creating synthetic training datasets, but in doing so they may transfer biases to downstream tasks. We focus…

cs.LG20231 cited

Learning Optimal Advantage from Preferences and Mistaking it for Reward

W. Bradley Knox, Stephane Hatgis-Kessell, Sigurdur Orn Adalgeirsson +4

We consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most re…