5 papers
Preserving Expert-Level Privacy in Offline Reinforcement Learning
Navodita Sharma, Vishnu Vinod, Abhradeep Thakurta +4
The offline reinforcement learning (RL) problem aims to learn an optimal policy from historical data collected by one or more behavioural policies (experts) by interacting with an…
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
Jiawei Huang, Bingcong Li, Christoph Dann +1
Sample efficiency is critical for online Reinforcement Learning from Human Feedback (RLHF). While existing works investigate sample-efficient online exploration strategies, the pot…
Mitigating Preference Hacking in Policy Optimization with Pessimism
Dhawal Gupta, Adam Fisch, Christoph Dann +1
This work tackles the problem of overoptimization in reinforcement learning from human feedback (RLHF), a prevalent technique for aligning models with human preferences. RLHF relie…
Design Considerations in Offline Preference-based RL
Alekh Agarwal, Christoph Dann, Teodor V. Marinov
Offline algorithms for Reinforcement Learning from Human Preferences (RLHF), which use only a fixed dataset of sampled responses given an input, and preference feedback among these…
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
Kaiwen Wang, Rahul Kidambi, Ryan Sullivan +17
Reward-based finetuning is crucial for aligning language policies with intended behaviors (e.g., creativity and safety). A key challenge is to develop steerable language models tha…