1 paper
Augusto Tagle, Javier Ruiz-del-Solar, Felipe Tobar
Offline reinforcement learning (RL) recovers the optimal policy I¨ given historical observations of an agent. In practice, I¨ is modeled as a weighted version of the agent's be…