1 paper
Wenhui Liu, Zhijian Wu, Jingchao Wang +2
Offline reinforcement learning seeks to derive improved policies entirely from historical data but often struggles with over-optimistic value estimates for out-of-distribution (OOD…