2 papers
cs.LG2024
RLHF and IIA: Perverse Incentives
Wanqiao Xu, Shi Dong, Xiuyuan Lu +3
Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independen…
cs.CL2023
Shattering the Agent-Environment Interface for Fine-Tuning Inclusive Language Models
Wanqiao Xu, Shi Dong, Dilip Arumugam +1
A centerpiece of the ever-popular reinforcement learning from human feedback (RLHF) approach to fine-tuning autoregressive language models is the explicit training of a reward mode…