1 citations · 1 across the 3 of their papers we have counts for
1 paper · 2 filters
Yuanying Cai, Chuheng Zhang, Li Zhao +6
We consider an offline reinforcement learning (RL) setting where the agent need to learn from a dataset collected by rolling out multiple behavior policies. There are two challenge…