activity
20182022
most citedPessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

24 citations · 80 across the 9 of their papers we have counts for

collaborators

13 papers

cs.LG202224 cited

Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement Learning

Chenjia Bai, Lingxiao Wang, Zhuoran Yang +4

Offline Reinforcement Learning (RL) aims to learn policies from previously collected datasets without exploring the environment. Directly applying off-policy algorithms to offline…

cs.LG202113 cited

Dynamic Bottleneck for Robust Self-Supervised Exploration

Chenjia Bai, Lingxiao Wang, Lei Han +4

Exploration methods based on pseudo-count of transitions or curiosity of dynamics have achieved promising results in solving reinforcement learning with sparse rewards. However, su…

cs.LG20212 cited

Adaptive Differentially Private Empirical Risk Minimization

Xiaoxia Wu, Lingxiao Wang, Irina Cristali +2

We propose an adaptive (stochastic) gradient perturbation method for differentially private empirical risk minimization. At each iteration, the random noise added to the gradient i…

cs.LG20213 cited

Permutation Invariant Policy Optimization for Mean-Field Multi-Agent Reinforcement Learning: A Principled Approach

Yan Li, Lingxiao Wang, Jiachen Yang +4

Multi-agent reinforcement learning (MARL) becomes more challenging in the presence of more agents, as the capacity of the joint state and action spaces grows exponentially in the n…

cs.LG20217 cited

Principled Exploration via Optimistic Bootstrapping and Backward Induction

Chenjia Bai, Lingxiao Wang, Lei Han +4

One principled approach for provably efficient exploration is incorporating the upper confidence bound (UCB) into the value function as a bonus. However, UCB is specified to deal w…

physics.soc-ph20205 cited

Machine learning spatio-temporal epidemiological model to evaluate Germany-county-level COVID-19 risk

Lingxiao Wang, Tian Xu, Till Hannes Stoecker +3

As the COVID-19 pandemic continues to ravage the world, it is of critical significance to provide a timely risk prediction of the COVID-19 in multi-level. To implement it and evalu…