6 papers
Corruption Robust Offline Reinforcement Learning with Human Feedback
Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban +2
We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedba…
Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks
Paulius Sasnauskas, YiÄit Yalın, Goran RadanoviÄ
We study the corruption-robustness of in-context reinforcement learning (ICRL), focusing on the Decision-Pretrained Transformer (DPT, Lee et al., 2023). To address the challenge of…
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban +2
We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset of…
AgenticRed: Evolving Agentic Systems for Red-Teaming
Jiayi Yuan, Jonathan Nöther, Natasha Jaques +1
While recent automated red-teaming methods show promise for systematically exposing model vulnerabilities, most existing approaches rely on human-specified workflows. This dependen…
Reinforcement Learning for Durable Algorithmic Recourse
Marina Ceccon, Alessandro Fabris, Goran RadanoviÄ +2
Algorithmic recourse seeks to provide individuals with actionable recommendations that increase their chances of receiving favorable outcomes from automated decision systems (e.g.,…
Independent Learning in Performative Markov Potential Games
Rilind Sahitaj, Paulius Sasnauskas, YiÄit Yalın +2
Performative Reinforcement Learning (PRL) refers to a scenario in which the deployed policy changes the reward and transition dynamics of the underlying environment. In this work,…