Variance Adjusted Actor Critic Algorithms
arXiv:1310.3697
Abstract
We present an actor-critic framework for MDPs where the objective is the variance-adjusted expected return. Our critic uses linear function approximation, and we extend the concept of compatible features to the variance-adjusted setting. We present an episodic actor-critic algorithm and show that it converges almost surely to a locally optimal point of the objective function.
References in corpus (1)
Cited by in corpus (10)
- Reward Constrained Policy Optimization
- Cumulative Prospect Theory Meets Reinforcement Learning: Prediction and Control
- Risk-Constrained Reinforcement Learning with Percentile Risk Criteria
- Continuous-Time Mean-Variance Portfolio Selection: A Reinforcement Learning Framework
- Constrained Reinforcement Learning Has Zero Duality Gap
- Risk-Sensitive Deep RL: Variance-Constrained Actor-Critic Provably Finds Globally Optimal Policy
- Directly Estimating the Variance of the λ-Return Using Temporal-Difference Methods
- Large scale continuous-time mean-variance portfolio allocation via reinforcement learning
- A Unified Off-Policy Evaluation Approach for General Value Function
- Model-Based Actor-Critic with Chance Constraint for Stochastic System