Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games
Tristan Maidment, JB Lanier, Chase McDonald +5
Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-…
cs.LG2026
Human-like autonomy emerges from self-play and a pinch of human data
Daphne Cornelisse, Julian Hunt, Zixu Zhang +4
Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data. It uses cheap, large-scale simulations to substitute expensive, lar…
cs.LG2025
Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
Samuel Sokota, Eugene Vinitsky, Hengyuan Hu +2
Few classical games have been regarded as such significant benchmarks of artificial intelligence as to have justified training costs in the millions of dollars. Among these, Strate…