1 paper
Thanh Vinh Vo, Young Lee, Haozhe Ma +2
Hidden confounders that influence both states and actions can bias policy learning in reinforcement learning (RL), leading to suboptimal or non-generalizable behavior. Most RL algo…