4 papers
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada +4
In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce un…
Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
Soichiro Nishimori, Shinri Okano, Keigo Habara +3
Riichi Mahjong is a multi-player, imperfect-information game characterized by stochasticity and high-dimensional state spaces. These attributes present a unique combination of chal…
A Simple, Solid, and Reproducible Baseline for Bridge Bidding AI
Haruka Kita, Sotetsu Koyamada, Yotaro Yamaguchi +1
Contract bridge, a cooperative game characterized by imperfect information and multi-agent dynamics, poses significant challenges and serves as a critical benchmark in artificial i…
A Batch Sequential Halving Algorithm without Performance Degradation
Sotetsu Koyamada, Soichiro Nishimori, Shin Ishii
In this paper, we investigate the problem of pure exploration in the context of multi-armed bandits, with a specific focus on scenarios where arms are pulled in fixed-size batches.…