activity
20242026
collaborators

9 papers

math.OC2026

Network Epidemic Control via Model Predictive Control: Extended Version

Mahtab Talaei, Alex Olshevsky, Laura F. White +1

Balancing the societal costs of non-pharmaceutical interventions with epidemic suppression requires adaptive feedback control. Rather than relying on state-dependent operational ca…

cs.LG2026

Bridging the Gap Between Average and Discounted TD Learning

Haoxing Tian, Zaiwei Chen, Ioannis Ch. Paschalidis +1

The analysis of Temporal Difference (TD) learning in the average-reward setting faces notable theoretical difficulties because the Bellman operator is not contractive with respect…

cs.LG2026

Data Deletion Can Help in Adaptive RL

Param Budhraja, Aditya Gangrade, Alex Olshevsky +1

Deploying reinforcement learning policies in the real world requires adapting to time-varying environments. We study this problem in the contextual Markov Decision Process (cMDP) f…

cs.LG2025

Geometric Re-Analysis of Classical MDP Solving Algorithms

Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1

We build on a recently introduced geometric interpretation of Markov Decision Processes (MDPs) to analyze classical MDP-solving algorithms: Value Iteration (VI) and Policy Iteratio…

cs.LG2025

MDP Geometry, Normalization and Reward Balancing Solvers

Arsenii Mustafin, Aleksei Pakharev, Alex Olshevsky +1

We present a new geometric interpretation of Markov Decision Processes (MDPs) with a natural normalization procedure that allows us to adjust the value function at each state witho…

cs.LG2025

Analysis of Value Iteration Through Absolute Probability Sequences

Arsenii Mustafin, Sebastien Colla, Alex Olshevsky +1

Value Iteration is a widely used algorithm for solving Markov Decision Processes (MDPs). While previous studies have extensively analyzed its convergence properties, they primarily…