5 papers
Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains
Rahul Singh, Siddharth Chandak, Eric Moulines +2
We present the first regret bound for classical online Q-learning in infinite-horizon discounted Markov decision processes (MDPs), without relying on optimism or bonus terms. We fi…
A Concentration Bound for TD(0) with Function Approximation
Siddharth Chandak, Vivek S. Borkar
We derive uniform all-time concentration bound of the type 'for all for some ' for TD(0) with linear function approximation. We work with online TD learning with…
Strong and weak quantitative estimates in slow-fast diffusions using filtering techniques
Sumith Reddy Anugu, Vivek S. Borkar
The behavior of slow-fast diffusions as the separation of scale diverges is a well-studied problem in the literature. In this short paper, we revisit this problem and obtain a new…
An Actor-Critic Algorithm with Function Approximation for Risk Sensitive Cost Markov Decision Processes
Soumyajit Guin, Vivek S. Borkar, Shalabh Bhatnagar
In this paper, we consider the risk-sensitive cost criterion with exponentiated costs for Markov decision processes and develop a model-free policy gradient algorithm in this setti…
Small noise limits of Markov chains and the PageRank
Vivek S Borkar, S Sowmya, Raghavendra Tripathi
We recall the classical formulation of PageRank as the stationary distribution of a singularly perturbed irreducible Markov chain that is not irreducible when the perturbation para…