1 paper
Xiaohong Chen, Yuling Jiao, Lican Kang +2
In offline RL, estimating the optimal action-value function Q∗ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental chal…