Generalized Emphatic Temporal Difference Learning: Bias-Variance Analysis
arXiv:1509.05172
Abstract
We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced \emph{emphatic temporal differences} (ETD) algorithm \citep{SuttonMW15}, which encompasses the original ETD(), as well as several other off-policy evaluation algorithms as special cases. We call this framework \ETD, where our introduced parameter controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying \ETD\ involves a contraction operator, allowing us to present the first asymptotic error bounds (bias) for \ETD. Our results show that the original ETD algorithm always involves a contraction operator, and its bias is bounded. Moreover, by controlling , our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.
arXiv admin note: text overlap with arXiv:1508.03411
References in corpus (4)
Cited by in corpus (10)
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Safe and Efficient Off-Policy Reinforcement Learning
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
- Unifying task specification in reinforcement learning
- Online Off-policy Prediction
- Generalized Off-Policy Actor-Critic
- Variance-Reduced Off-Policy Memory-Efficient Policy Search
- Truncated Emphatic Temporal Difference Methods for Prediction and Control
- A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
- An Empirical Comparison of Off-policy Prediction Learning Algorithms in the Four Rooms Environment