activity
20212025
most citedOnline Policy Optimization for Robust MDP

1 citations · 3 across the 14 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2025

Learning Multi-Timescale Interventions under Safety and Resource Constraints

David Mguni, Wanrong Yang, Jing Dong +5

Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects th…

cs.LG2023

DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning

Canzhe Zhao, Yanjie Ze, Jing Dong +2

Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating…

cs.LG2023

A Batch-to-Online Transformation under Random-Order Model

Jing Dong, Yuichi Yoshida

We introduce a transformation framework that can be utilized to develop online algorithms with low -approximate regret in the random-order model from offline approximation algor…

cs.LG2022★ 1 cited

Online Policy Optimization for Robust MDP

Jing Dong, Jingwei Li, Baoxiang Wang +1

Reinforcement learning (RL) has exceeded human performance in many synthetic settings such as video games and Go. However, real-world deployment of end-to-end RL models is less com…

cs.LG2022★ 1 cited

Algorithms and Theory for Supervised Gradual Domain Adaptation

Jing Dong, Shiji Zhou, Baoxiang Wang +1

The phenomenon of data distribution evolving over time has been observed in a range of applications, calling the needs of adaptive learning algorithms. We thus study the problem of…

cs.LG2022★ 1 cited

Provably Efficient Convergence of Primal-Dual Actor-Critic with Nonlinear Function Approximation

Jing Dong, Li Shen, Yinggan Xu +1

We study the convergence of the actor-critic algorithm with nonlinear function approximation under a nonconvex-nonconcave primal-dual formulation. Stochastic gradient descent ascen…