1 paper · 1 filter
Tsunehiko Tanaka, Kenshi Abe, Kaito Ariu +2
Traditional approaches in offline reinforcement learning aim to learn the optimal policy that maximizes the cumulative reward, also known as return. It is increasingly important to…