3 papers
cs.LG2026
Distributionally Robust Reinforcement Learning with Human Feedback
Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic
Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). However, existing RLHF methods are non-rob…
cs.LG2026
Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks
Paulius Sasnauskas, YiÄit Yalın, Goran RadanoviÄ
We study the corruption-robustness of in-context reinforcement learning (ICRL), focusing on the Decision-Pretrained Transformer (DPT, Lee et al., 2023). To address the challenge of…
cs.LG2025
Independent Learning in Performative Markov Potential Games
Rilind Sahitaj, Paulius Sasnauskas, YiÄit Yalın +2
Performative Reinforcement Learning (PRL) refers to a scenario in which the deployed policy changes the reward and transition dynamics of the underlying environment. In this work,…