3 papers
cs.LG2025
Mitigating Preference Hacking in Policy Optimization with Pessimism
Dhawal Gupta, Adam Fisch, Christoph Dann +1
This work tackles the problem of overoptimization in reinforcement learning from human feedback (RLHF), a prevalent technique for aligning models with human preferences. RLHF relie…
cs.LG2024
ICU-Sepsis: A Benchmark MDP Built from Real Medical Data
Kartik Choudhary, Dhawal Gupta, Philip S. Thomas
We present ICU-Sepsis, an environment that can be used in benchmarks for evaluating reinforcement learning (RL) algorithms. Sepsis management is a complex task that has been an imp…
cs.AI2024
A Safe Exploration Strategy for Model-free Task Adaptation in Safety-constrained Grid Environments
Erfan Entezami, Mahsa Sahebdel, Dhawal Gupta
Training a model-free reinforcement learning agent requires allowing the agent to sufficiently explore the environment to search for an optimal policy. In safety-constrained enviro…