4 papers
A Descriptive and Normative Theory of Human Beliefs in RLHF
Sylee Dandekar, Shripad Deshmukh, Frank Chiu +2
Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values. In this work, we propose that human belie…
Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach
Kihyun Kim, Shripad Deshmukh, Nikos Vlassis +1
Inverse reinforcement learning (IRL) typically assumes demonstrations from a single optimal demonstrator, but in many applications data come from multiple imperfect demonstrators w…
COSAC: Counterfactual Credit Assignment in Sequential Cooperative Teams
Shripad Deshmukh, Jayakumar Subramanian, Raghavendra Addanki +1
In cooperative teams where agents act in a fixed order and share a single team-level reward (multi-agent language systems, sequential robotic tasks), per-agent credit assignment is…
Evaluation-Aware Reinforcement Learning
Shripad Vilasrao Deshmukh, Will Schwarzer, Scott Niekum
Policy evaluation is a core component of many reinforcement learning (RL) algorithms and a critical tool for ensuring safe deployment of RL policies. However, existing policy evalu…