5 papers
Ergodicity in reinforcement learning
Dominik Baumann, Erfaun Noorani, Arsenii Mustafin +5
In reinforcement learning, we typically aim to optimize the expected value of the sum of rewards an agent collects over a trajectory. However, if the process generating these rewar…
Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases
Arsenii Mustafin, Xinyi Sheng, Dominik Baumann
While Value Iteration (VI) is one of the most fundamental algorithms in Reinforcement Learning, its theoretical convergence guarantees still exhibit a persistent mismatch with empi…
Geometric Re-Analysis of Classical MDP Solving Algorithms
Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1
We build on a recently introduced geometric interpretation of Markov Decision Processes (MDPs) to analyze classical MDP-solving algorithms: Value Iteration (VI) and Policy Iteratio…
MDP Geometry, Normalization and Reward Balancing Solvers
Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1
We present a new geometric interpretation of Markov Decision Processes (MDPs) with a natural normalization procedure that allows us to adjust the value function at each state witho…
Analysis of Value Iteration Through Absolute Probability Sequences
Arsenii Mustafin, Sebastien Colla, Alex Olshevsky +1
Value Iteration is a widely used algorithm for solving Markov Decision Processes (MDPs). While previous studies have extensively analyzed its convergence properties, they primarily…