activity
20152026
most citedFinite-Time Analysis of Distributed TD(0) with Linear Function Approximation for Multi-Agent Reinforcement Learning

50 citations · 117 across the 38 of their papers we have counts for

collaborators
Showing 2021Show all

10 papers · 1 filter

cs.LG2021

Stationary Behavior of Constant Stepsize SGD Type Algorithms: An Asymptotic Characterization

Zaiwei Chen, Shancong Mou, Siva Theja Maguluri

Stochastic approximation (SA) and stochastic gradient descent (SGD) algorithms are work-horses for modern machine learning algorithms. Their constant stepsize variants are preferre…

math.PR2021

A Heavy Traffic Theory of Matching Queues

Sushil Mahavir Varma, Siva Theja Maguluri

Motivated by emerging applications in online matching platforms and marketplaces, we study a matching queue. Customers and servers that arrive in a matching queue depart as soon as…

cs.NI2021★ 2 cited

Transportation Polytope and its Applications in Parallel Server Systems

Sushil Mahavir Varma, Siva Theja Maguluri

A parallel server system is a stochastic processing network with applications in manufacturing, supply chain, ride-hailing, call centers, etc. Heterogeneous customers arrive in the…

cs.LG2021★ 4 cited

Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman Operators

Zaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai +1

In temporal difference (TD) learning, off-policy sampling is known to be more practical than on-policy sampling, and by decoupling learning from data collection, it enables data re…

math.OC2021★ 3 cited

Optimal Pricing in Multi Server Systems

Ashok Krishnan K. S, Chandramani Singh, Siva Theja Maguluri +1

We study optimal service pricing in server farms where customers arrive according to a renewal process and have independent and identical () exponential service times and $…

cs.LG2021

On the Linear convergence of Natural Policy Gradient Algorithm

Sajad Khodadadian, Prakirt Raj Jhunjhunwala, Sushil Mahavir Varma +1

Markov Decision Processes are classically solved using Value Iteration and Policy Iteration algorithms. Recent interest in Reinforcement Learning has motivated the study of methods…