6 papers
Preserving Expert-Level Privacy in Offline Reinforcement Learning
Navodita Sharma, Vishnu Vinod, Abhradeep Thakurta +4
The offline reinforcement learning (RL) problem aims to learn an optimal policy from historical data collected by one or more behavioural policies (experts) by interacting with an…
Mitigating Preference Hacking in Policy Optimization with Pessimism
Dhawal Gupta, Adam Fisch, Christoph Dann +1
This work tackles the problem of overoptimization in reinforcement learning from human feedback (RLHF), a prevalent technique for aligning models with human preferences. RLHF relie…
Design Considerations in Offline Preference-based RL
Alekh Agarwal, Christoph Dann, Teodor V. Marinov
Offline algorithms for Reinforcement Learning from Human Preferences (RLHF), which use only a fixed dataset of sampled responses given an input, and preference feedback among these…
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
Chenlu Ye, Yujia Jin, Alekh Agarwal +1
Typical contextual bandit algorithms assume that the rewards at each round lie in some fixed range , and their regret scales polynomially with this reward range . Howeve…
Peer Reviews of Peer Reviews: A Randomized Controlled Trial and Other Experiments
Alexander Goldberg, Ivan Stelmakh, Kyunghyun Cho +4
Is it possible to reliably evaluate the quality of peer reviews? We study this question driven by two primary motivations -- incentivizing high-quality reviewing using assessed qua…
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
Kaiwen Wang, Rahul Kidambi, Ryan Sullivan +17
Reward-based finetuning is crucial for aligning language policies with intended behaviors (e.g., creativity and safety). A key challenge is to develop steerable language models tha…