activity
20242026
collaborators

5 papers

cs.LG2026

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

Tristan Maidment, JB Lanier, Chase McDonald +5

Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-…

cs.LG2026

Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games

JB Lanier, Nathan Monette, Pierre Baldi +1

Finding approximate equilibria for large-scale imperfect-information competitive games such as StarCraft, Dota, and CounterStrike remains computationally infeasible due to sparse r…

cs.LG2026

Model-Based Reinforcement Learning under Random Observation Delays

Armin Karamzade, Kyungmin Kim, JB Lanier +2

Delays frequently occur in real-world environments, yet standard reinforcement learning (RL) algorithms often assume instantaneous perception of the environment. We study random se…

cs.LG2025

Adapting World Models with Latent-State Dynamics Residuals

JB Lanier, Kyungmin Kim, Armin Karamzade +5

Simulation-to-reality reinforcement learning (RL) faces the critical challenge of reconciling discrepancies between simulated and real-world dynamics, which can severely degrade ag…

cs.LG2024

Realizable Continuous-Space Shields for Safe Reinforcement Learning

Kyungmin Kim, Davide Corsi, Andoni Rodriguez +5

While Deep Reinforcement Learning (DRL) has achieved remarkable success across various domains, it remains vulnerable to occasional catastrophic failures without additional safegua…