1 paper · 1 filter
Tristan Maidment, JB Lanier, Chase McDonald +5
Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-…