46 citations · 52 across the 2 of their papers we have counts for
3 papers · 1 filter
Jackpot! Alignment as a Maximal Lottery
Roberto-Rafael Maura-Rivero, Marc Lanctot, Francesco Visin +1
Reinforcement Learning from Human Feedback (RLHF), the standard for aligning Large Language Models (LLMs) with human values, is known to fail to satisfy properties that are intuiti…
Neural Population Learning beyond Symmetric Zero-sum Games
Siqi Liu, Luke Marris, Marc Lanctot +3
We study computationally efficient methods for finding equilibria in n-player general-sum games, specifically ones that afford complex visuomotor skills. We show how existing metho…
Monte Carlo Tree Search with Heuristic Evaluations using Implicit Minimax Backups
Marc Lanctot, Mark H. M. Winands, Tom Pepels +1
Monte Carlo Tree Search (MCTS) has improved the performance of game engines in domains such as Go, Hex, and general game playing. MCTS has been shown to outperform classic alpha-be…