5 papers
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games
Tristan Maidment, JB Lanier, Chase McDonald +5
Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-…
Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games
JB Lanier, Nathan Monette, Pierre Baldi +1
Finding approximate equilibria for large-scale imperfect-information competitive games such as StarCraft, Dota, and CounterStrike remains computationally infeasible due to sparse r…
Model-Based Reinforcement Learning under Random Observation Delays
Armin Karamzade, Kyungmin Kim, JB Lanier +2
Delays frequently occur in real-world environments, yet standard reinforcement learning (RL) algorithms often assume instantaneous perception of the environment. We study random se…
Adapting World Models with Latent-State Dynamics Residuals
JB Lanier, Kyungmin Kim, Armin Karamzade +5
Simulation-to-reality reinforcement learning (RL) faces the critical challenge of reconciling discrepancies between simulated and real-world dynamics, which can severely degrade ag…
Realizable Continuous-Space Shields for Safe Reinforcement Learning
Kyungmin Kim, Davide Corsi, Andoni Rodriguez +5
While Deep Reinforcement Learning (DRL) has achieved remarkable success across various domains, it remains vulnerable to occasional catastrophic failures without additional safegua…