most citedFinite-Time Analysis of Distributed TD(0) with Linear Function Approximation for Multi-Agent Reinforcement Learning

50 citations · 63 across the 5 of their papers we have counts for

collaborators

8 papers

cs.DC20203 cited

Byzantine Fault-Tolerance in Decentralized Optimization under Minimal Redundancy

Nirupam Gupta, Thinh T. Doan, Nitin H. Vaidya

This paper considers the problem of Byzantine fault-tolerance in multi-agent decentralized optimization. In this problem, each agent has a local cost function. The goal of a decent…

cs.LG20206 cited

Local Stochastic Approximation: A Unified View of Federated Learning and Distributed Multi-Task Reinforcement Learning Algorithms

Thinh T. Doan

Motivated by broad applications in reinforcement learning and federated learning, we study local stochastic approximation over a network of agents, where their goal is to find the…

math.OC2020

Finite-Time Analysis of Stochastic Gradient Descent under Markov Randomness

Thinh T. Doan, Lam M. Nguyen, Nhan H. Pham +1

Motivated by broad applications in reinforcement learning and machine learning, this paper considers the popular stochastic gradient descent (SGD) when the gradients of the underly…

cs.LG20202 cited

Finite-Time Analysis and Restarting Scheme for Linear Two-Time-Scale Stochastic Approximation

Thinh T. Doan

Motivated by their broad applications in reinforcement learning, we study the linear two-time-scale stochastic approximation, an iterative method using two different step sizes for…

math.OC20192 cited

Finite-Time Performance of Distributed Two-Time-Scale Stochastic Approximation

Thinh T. Doan, Justin Romberg

Two-time-scale stochastic approximation is a popular iterative method for finding the solution of a system of two equations. Such methods have found broad applications in many area…

cs.RO2019

A Reinforcement Learning Framework for Sequencing Multi-Robot Behaviors

Pietro Pierpaoli, Thinh T. Doan, Justin Romberg +1

Given a list of behaviors and associated parameterized controllers for solving different individual tasks, we study the problem of selecting an optimal sequence of coordinated beha…