1 paper
Songyuan Zhang, Oswin So, H. M. Sabbir Ahmad +4
Offline reinforcement learning (RL) aims to learn the optimal policy from a fixed dataset generated by behavior policies without additional environment interactions. One common cha…