1 paper
Haoyu Wang, Jingcheng Wang, Shunyu Wu +1
Offline reinforcement learning (RL) can fit strong value functions from fixed datasets, yet reliable deployment still hinges on the action selection interface used to query them. W…