1 citations · 1 across the 4 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.LG2026
What Emerges and What Breaks in Self-Play Driving
Laur Sisask, Ardi Tampuu, Tambet Matiisen
Training autonomous driving policies through pure self-play has recently shown promising results. Following Gigaflow and Puffer- Drive, we train driving policies in a similar self-…
cs.LG2026
Reproducing AlphaZero on Tablut: Self-Play RL for an Asymmetric Board Game
Tõnis Lees, Tambet Matiisen
This work investigates the adaptation of the AlphaZero reinforcement learning algorithm to Tablut, an asymmetric historical board game featuring unequal piece counts and distinct p…