3 papers
cs.LG2026
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada +4
In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce un…
cs.AI2024
A Simple, Solid, and Reproducible Baseline for Bridge Bidding AI
Haruka Kita, Sotetsu Koyamada, Yotaro Yamaguchi +1
Contract bridge, a cooperative game characterized by imperfect information and multi-agent dynamics, poses significant challenges and serves as a critical benchmark in artificial i…
stat.ML2024
A Batch Sequential Halving Algorithm without Performance Degradation
Sotetsu Koyamada, Soichiro Nishimori, Shin Ishii
In this paper, we investigate the problem of pure exploration in the context of multi-armed bandits, with a specific focus on scenarios where arms are pulled in fixed-size batches.…