1 paper
Renos Zabounidis, Roy Siegelmann, Mohamad Qadri +3
In reinforcement learning environments with state-dependent action validity, action masking consistently outperforms penalty-based handling of invalid actions, yet existing theory…