2 papers
cs.LG2026
Online Policy Evaluation for MDPs with Dynamic UBSR Measures
Weikai Wang, Erick Delage
Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restr…
cs.LG2025
Planning and Learning in Average Risk-aware MDPs
Weikai Wang, Erick Delage
For continuing tasks, average cost Markov decision processes have well-documented value and can be solved using efficient algorithms. However, it explicitly assumes that the agent…