9 papers
Network Epidemic Control via Model Predictive Control: Extended Version
Mahtab Talaei, Alex Olshevsky, Laura F. White +1
Balancing the societal costs of non-pharmaceutical interventions with epidemic suppression requires adaptive feedback control. Rather than relying on state-dependent operational ca…
Bridging the Gap Between Average and Discounted TD Learning
Haoxing Tian, Zaiwei Chen, Ioannis Ch. Paschalidis +1
The analysis of Temporal Difference (TD) learning in the average-reward setting faces notable theoretical difficulties because the Bellman operator is not contractive with respect…
Data Deletion Can Help in Adaptive RL
Param Budhraja, Aditya Gangrade, Alex Olshevsky +1
Deploying reinforcement learning policies in the real world requires adapting to time-varying environments. We study this problem in the contextual Markov Decision Process (cMDP) f…
Geometric Re-Analysis of Classical MDP Solving Algorithms
Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1
We build on a recently introduced geometric interpretation of Markov Decision Processes (MDPs) to analyze classical MDP-solving algorithms: Value Iteration (VI) and Policy Iteratio…
MDP Geometry, Normalization and Reward Balancing Solvers
Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1
We present a new geometric interpretation of Markov Decision Processes (MDPs) with a natural normalization procedure that allows us to adjust the value function at each state witho…
Analysis of Value Iteration Through Absolute Probability Sequences
Arsenii Mustafin, Sebastien Colla, Alex Olshevsky +1
Value Iteration is a widely used algorithm for solving Markov Decision Processes (MDPs). While previous studies have extensively analyzed its convergence properties, they primarily…