4 papers
Reevaluating Policy Gradient Methods for Imperfect-Information Games
Max Rudolph, Nathan Lichtle, Sobhan Mohammadpour +6
In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed nu…
Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
Samuel Sokota, Eugene Vinitsky, Hengyuan Hu +2
Few classical games have been regarded as such significant benchmarks of artificial intelligence as to have justified training costs in the millions of dollars. Among these, Strate…
Computing Low-Entropy Couplings for Large-Support Distributions
Samuel Sokota, Dylan Sam, Christian Schroeder de Witt +3
Minimum-entropy coupling (MEC) -- the process of finding a joint distribution with minimum entropy for given marginals -- has applications in areas such as causality and steganogra…
The Update-Equivalence Framework for Decision-Time Planning
Samuel Sokota, Gabriele Farina, David J. Wu +4
The process of revising (or constructing) a policy at execution time -- known as decision-time planning -- has been key to achieving superhuman performance in perfect-information g…