paper

Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ

arXiv:2402.11877

Abstract

Reinforcement learning has witnessed significant advancements, particularly with the emergence of model-based approaches. Among these, -learning has proven to be a powerful algorithm in model-free settings. However, the extension of -learning to a model-based framework remains relatively unexplored. In this paper, we investigate the sample complexity of -learning when integrated with a model-based approach. The proposed algorihtms learns both the model and Q-value in an online manner. We demonstrate a near-optimal sample complexity result within a broad range of step sizes.

Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ · wovepaper