19 citations · 56 across the 10 of their papers we have counts for
8 papers · 1 filter
Human-AI Coordination via Human-Regularized Search and Learning
Hengyuan Hu, David J Wu, Adam Lerer +2
We consider the problem of making AI agents that collaborate well with humans in partially observable fully cooperative environments given datasets of human behavior. Inspired by p…
Scalable Online Planning via Reinforcement Learning Fine-Tuning
Arnaud Fickinger, Hengyuan Hu, Brandon Amos +2
Lookahead search has been a critical component of recent AI successes, such as in the games of chess, go, and poker. However, the search methods used in these games, and in many ot…
Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings
Hengyuan Hu, Adam Lerer, Noam Brown +1
Search is an important tool for computing effective policies in single- and multi-agent environments, and has been crucial for achieving superhuman performance in several benchmark…
Off-Belief Learning
Hengyuan Hu, Adam Lerer, Brandon Cui +4
The standard problem setting in Dec-POMDPs is self-play, where the goal is to find a set of policies that play optimally together. Policies learned through self-play may adopt arbi…
Unlocking the Potential of Deep Counterfactual Value Networks
Ryan Zarick, Bryan Pellegrino, Noam Brown +1
Deep counterfactual value networks combined with continual resolving provide a way to conduct depth-limited search in imperfect-information games. However, since their introduction…
Improving Policies via Search in Cooperative Partially Observable Games
Adam Lerer, Hengyuan Hu, Jakob Foerster +1
Recent superhuman results in games have largely been achieved in a variety of zero-sum settings, such as Go and Poker, in which agents need to compete against others. However, just…