collaborators

5 papers

cs.LG2026

Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains

Rahul Singh, Siddharth Chandak, Eric Moulines +2

We present the first regret bound for classical online Q-learning in infinite-horizon discounted Markov decision processes (MDPs), without relying on optimism or bonus terms. We fi…

cs.LG2026

A Concentration Bound for TD(0) with Function Approximation

Siddharth Chandak, Vivek S. Borkar

We derive uniform all-time concentration bound of the type 'for all for some ' for TD(0) with linear function approximation. We work with online TD learning with…

math.OC2025

Strong and weak quantitative estimates in slow-fast diffusions using filtering techniques

Sumith Reddy Anugu, Vivek S. Borkar

The behavior of slow-fast diffusions as the separation of scale diverges is a well-studied problem in the literature. In this short paper, we revisit this problem and obtain a new…

cs.LG2025

An Actor-Critic Algorithm with Function Approximation for Risk Sensitive Cost Markov Decision Processes

Soumyajit Guin, Vivek S. Borkar, Shalabh Bhatnagar

In this paper, we consider the risk-sensitive cost criterion with exponentiated costs for Markov decision processes and develop a model-free policy gradient algorithm in this setti…

math.PR2025

Small noise limits of Markov chains and the PageRank

Vivek S Borkar, S Sowmya, Raghavendra Tripathi

We recall the classical formulation of PageRank as the stationary distribution of a singularly perturbed irreducible Markov chain that is not irreducible when the perturbation para…