3 papers
cs.LG2025
Learning a Pessimistic Reward Model in RLHF
Yinglun Xu, Hangoo Kang, Tarun Suresh +2
This work proposes `PET', a novel pessimistic reward fine-tuning method, to learn a pessimistic reward model robust against reward hacking in offline reinforcement learning from hu…
cs.LG2023
On the Robustness of Epoch-Greedy in Multi-Agent Contextual Bandit Mechanisms
Yinglun Xu, Bhuvesh Kumar, Jacob Abernethy
Efficient learning in multi-armed bandit mechanisms such as pay-per-click (PPC) auctions typically involves three challenges: 1) inducing truthful bidding behavior (incentives), 2)…
cs.LG2023
Black-Box Targeted Reward Poisoning Attack Against Online Deep Reinforcement Learning
Yinglun Xu, Gagandeep Singh
We propose the first black-box targeted attack against online deep reinforcement learning through reward poisoning during training time. Our attack is applicable to general environ…