1 paper
Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada +4
In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce un…