22 citations · 32 across the 2 of their papers we have counts for
2 papers
cs.LG2022★ 22 cited
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Alexander Pan, Kush Bhatia, Jacob Steinhardt
Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied. To understand how reward hacking ar…
eess.SY2021★ 10 cited
Improving Robustness of Reinforcement Learning for Power System Control with Adversarial Training
Alexander Pan, Yongkyun Lee, Huan Zhang +2
Due to the proliferation of renewable energy and its intrinsic intermittency and stochasticity, current power systems face severe operational challenges. Data-driven decision-makin…