1 paper
Dan Elbaz, Oren Salzman
Offline Reinforcement Learning (RL) algorithms learn a policy using a fixed training dataset, which is then deployed online to interact with the environment and make decisions. Tra…