19 citations · 68 across the 15 of their papers we have counts for
3 papers · 2 filters
Scalable Online Planning via Reinforcement Learning Fine-Tuning
Arnaud Fickinger, Hengyuan Hu, Brandon Amos +2
Lookahead search has been a critical component of recent AI successes, such as in the games of chess, go, and poker. However, the search methods used in these games, and in many ot…
Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings
Hengyuan Hu, Adam Lerer, Noam Brown +1
Search is an important tool for computing effective policies in single- and multi-agent environments, and has been crucial for achieving superhuman performance in several benchmark…
Off-Belief Learning
Hengyuan Hu, Adam Lerer, Brandon Cui +4
The standard problem setting in Dec-POMDPs is self-play, where the goal is to find a set of policies that play optimally together. Policies learned through self-play may adopt arbi…