8 citations · 15 across the 5 of their papers we have counts for
1 paper · 1 filter
Yuanying Cai, Chuheng Zhang, Li Zhao +6
We consider an offline reinforcement learning (RL) setting where the agent need to learn from a dataset collected by rolling out multiple behavior policies. There are two challenge…