Publications (21)
Look-ahead Search on Top of Policy Networks in Imperfect Information Games
Ondrej Kubicek, Neil Burch, Viliam Lisy
Search in test time is often used to improve the performance of reinforcement learning algorithms. Performing theoretically sound search in fully adversarial two-player games with…
Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning
Julien Perolat, Bart de Vylder, Daniel Hennes +31
We introduce DeepNash, an autonomous agent capable of learning to play the imperfect information game Stratego from scratch, up to a human expert level. Stratego is one of the few…
Sound Algorithms in Imperfect Information Games
Michal Å ustr, Martin Schmid, Matej MoravÄÃk +3
Search has played a fundamental role in computer game research since the very beginning. And while online search has been commonly used in perfect information games such as Chess a…
Revisiting CFR+ and Alternating Updates
Neil Burch, Matej Moravcik, Martin Schmid
The CFR+ algorithm for solving imperfect information games is a variant of the popular CFR algorithm, with faster empirical performance on a range of problems. It was introduced wi…
Bayes' Bluff: Opponent Modelling in Poker
Finnegan Southey, Michael P. Bowling, Bryce Larson +4
Poker is a challenging problem for artificial intelligence, with non-deterministic dynamics, partial observability, and the added difficulty of unknown adversaries. Modelling all o…
DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
Matej MoravÄÃk, Martin Schmid, Neil Burch +7
Artificial intelligence has seen several breakthroughs in recent years, with games often serving as milestones. A common feature of these games is that players have perfect informa…
Predicting the Performance of IDA* using Conditional Distributions
Uzi Zahavi, Ariel Felner, Neil Burch +1
Korf, Reid, and Edelkamp introduced a formula to predict the number of nodes IDA* will expand on a single iteration for a given consistent heuristic, and experimentally demonstrate…
Population-based Evaluation in Repeated Rock-Paper-Scissors as a Benchmark for Multiagent Reinforcement Learning
Marc Lanctot, John Schultz, Neil Burch +4
Progress in fields of machine learning and adversarial planning has benefited significantly from benchmark domains, from checkers and the classic UCI data sets to Go and Diplomacy.…
Bayesian Action Decoder for Deep Multi-Agent Reinforcement Learning
Jakob N. Foerster, Francis Song, Edward Hughes +5
When observing the actions of others, humans make inferences about why they acted as they did, and what this implies about the world; humans also use the fact that their actions wi…
Coachable agents for interactive gameplay
Roberto Capobianco, Harm van Seijen, Nolan D. Bard +38
Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation m…
Human-Agent Cooperation in Bridge Bidding
Edward Lockhart, Neil Burch, Nolan Bard +4
We introduce a human-compatible reinforcement-learning approach to a cooperative game, making use of a third-party hand-coded human-compatible bot to generate initial training data…
Solving Imperfect Information Games Using Decomposition
Neil Burch, Michael Johanson, Michael Bowling
Decomposition, i.e. independently analyzing possible subgames, has proven to be an essential principle for effective decision-making in perfect information games. However, in imper…
From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization
Julien Perolat, Remi Munos, Jean-Baptiste Lespiau +10
In this paper we investigate the Follow the Regularized Leader dynamics in sequential imperfect information games (IIG). We generalize existing results of Poincaré recurrence from…
The Hanabi Challenge: A New Frontier for AI Research
Nolan Bard, Jakob N. Foerster, Sarath Chandar +12
From the early days of computing, games have been important testbeds for studying how well machines can do sophisticated decision making. In recent years, machine learning has made…
Variance Reduction in Monte Carlo Counterfactual Regret Minimization (VR-MCCFR) for Extensive Form Games using Baselines
Martin Schmid, Neil Burch, Marc Lanctot +3
Learning strategies for imperfect information games from samples of interaction is a challenging problem. A common method for this setting, Monte Carlo Counterfactual Regret Minimi…
AIVAT: A New Variance Reduction Technique for Agent Evaluation in Imperfect Information Games
Neil Burch, Martin Schmid, Matej MoravÄÃk +1
Evaluating agent performance when outcomes are stochastic and agents use randomized strategies can be challenging when there is limited data available. The variance of sampled outc…
Student of Games: A unified learning algorithm for both perfect and imperfect information games
Martin Schmid, Matej Moravcik, Neil Burch +10
Games have a long history as benchmarks for progress in artificial intelligence. Approaches using search and learning produced strong performance across many perfect information ga…
Rethinking Formal Models of Partially Observable Multiagent Decision Making
VojtÄch KovaÅÃk, Martin Schmid, Neil Burch +2
Multiagent decision-making in partially observable environments is usually modelled as either an extensive-form game (EFG) in game theory or a partially observable stochastic game…
Approximate exploitability: Learning a best response in large games
Finbarr Timbers, Nolan Bard, Edward Lockhart +6
Researchers have demonstrated that neural networks are vulnerable to adversarial examples and subtle environment changes, both of which one can view as a form of distribution shift…
Solving Common-Payoff Games with Approximate Policy Iteration
Samuel Sokota, Edward Lockhart, Finbarr Timbers +6
For artificially intelligent learning systems to have widespread applicability in real-world settings, it is important that they be able to operate decentrally. Unfortunately, dece…
No-Regret Learning in Extensive-Form Games with Imperfect Recall
Marc Lanctot, Richard Gibson, Neil Burch +2
Counterfactual Regret Minimization (CFR) is an efficient no-regret learning algorithm for decision problems modeled as extensive games. CFR's regret bounds depend on the requiremen…