16.2k citations · 16.2k across the 10 of their papers we have counts for
10 papers · 1 filter
Human-AI Coordination via Human-Regularized Search and Learning
Hengyuan Hu, David J Wu, Adam Lerer +2
We consider the problem of making AI agents that collaborate well with humans in partially observable fully cooperative environments given datasets of human behavior. Inspired by p…
Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings
Hengyuan Hu, Adam Lerer, Noam Brown +1
Search is an important tool for computing effective policies in single- and multi-agent environments, and has been crucial for achieving superhuman performance in several benchmark…
Off-Belief Learning
Hengyuan Hu, Adam Lerer, Brandon Cui +4
The standard problem setting in Dec-POMDPs is self-play, where the goal is to find a set of policies that play optimally together. Policies learned through self-play may adopt arbi…
Improving Policies via Search in Cooperative Partially Observable Games
Adam Lerer, Hengyuan Hu, Jakob Foerster +1
Recent superhuman results in games have largely been achieved in a variety of zero-sum settings, such as Go and Poker, in which agents need to compete against others. However, just…
Deep Counterfactual Regret Minimization
Noam Brown, Adam Lerer, Sam Gross +1
Counterfactual Regret Minimization (CFR) is the leading framework for solving large imperfect-information games. It converges to an equilibrium by iteratively traversing the game t…
Learning Existing Social Conventions via Observationally Augmented Self-Play
Adam Lerer, Alexander Peysakhovich
In order for artificial agents to coordinate effectively with people, they must act consistently with existing conventions (e.g. how to navigate in traffic, which language to speak…