collaborators

6 papers

cs.CR2025

Preserving Expert-Level Privacy in Offline Reinforcement Learning

Navodita Sharma, Vishnu Vinod, Abhradeep Thakurta +4

The offline reinforcement learning (RL) problem aims to learn an optimal policy from historical data collected by one or more behavioural policies (experts) by interacting with an…

cs.LG2025

Mitigating Preference Hacking in Policy Optimization with Pessimism

Dhawal Gupta, Adam Fisch, Christoph Dann +1

This work tackles the problem of overoptimization in reinforcement learning from human feedback (RLHF), a prevalent technique for aligning models with human preferences. RLHF relie…

cs.LG2025

Design Considerations in Offline Preference-based RL

Alekh Agarwal, Christoph Dann, Teodor V. Marinov

Offline algorithms for Reinforcement Learning from Human Preferences (RLHF), which use only a fixed dataset of sampled responses given an input, and preference feedback among these…

stat.ML2025

Catoni Contextual Bandits are Robust to Heavy-tailed Rewards

Chenlu Ye, Yujia Jin, Alekh Agarwal +1

Typical contextual bandit algorithms assume that the rewards at each round lie in some fixed range , and their regret scales polynomially with this reward range . Howeve…

cs.DL2024

Peer Reviews of Peer Reviews: A Randomized Controlled Trial and Other Experiments

Alexander Goldberg, Ivan Stelmakh, Kyunghyun Cho +4

Is it possible to reliably evaluate the quality of peer reviews? We study this question driven by two primary motivations -- incentivizing high-quality reviewing using assessed qua…

cs.LG2024

Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning

Kaiwen Wang, Rahul Kidambi, Ryan Sullivan +17

Reward-based finetuning is crucial for aligning language policies with intended behaviors (e.g., creativity and safety). A key challenge is to develop steerable language models tha…