activity
20182022
most citedSMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving

103 citations · 273 across the 23 of their papers we have counts for

collaborators
Showing 2021Show all

12 papers · 1 filter

cs.AI2021

Measuring the Non-Transitivity in Chess

Ricky Sanjaya, Jun Wang, Yaodong Yang

It has long been believed that Chess is the \emph{Drosophila} of Artificial Intelligence (AI). Studying Chess can productively provide valid knowledge about complex systems. Althou…

cs.LG20212 cited

Revisiting the Characteristics of Stochastic Gradient Noise and Dynamics

Yixin Wu, Rui Luo, Chen Zhang +2

In this paper, we characterize the noise of stochastic gradients and analyze the noise-induced dynamics during training deep neural networks by gradient-based optimizers. Specifica…

stat.ML2021

Viscos Flows: Variational Schur Conditional Sampling With Normalizing Flows

Vincent Moens, Aivar Sootla, Haitham Bou Ammar +1

We present a method for conditional sampling for pre-trained normalizing flows when only part of an observation is available. We derive a lower bound to the conditioning variable l…

cs.MA20214 cited

A Game-Theoretic Approach to Multi-Agent Trust Region Optimization

Ying Wen, Hui Chen, Yaodong Yang +4

Trust region methods are widely applied in single-agent reinforcement learning problems due to their monotonic performance-improvement guarantee at every iteration. Nonetheless, wh…

cs.MA202125 cited

MALib: A Parallel Framework for Population-based Multi-agent Reinforcement Learning

Ming Zhou, Ziyu Wan, Hanjing Wang +6

Population-based multi-agent reinforcement learning (PB-MARL) refers to the series of methods nested with reinforcement learning (RL) algorithms, which produces a self-generated se…

cs.LG202111 cited

High-Dimensional Bayesian Optimisation with Variational Autoencoders and Deep Metric Learning

Antoine Grosnit, Rasul Tutunov, Alexandre Max Maraval +9

We introduce a method combining variational autoencoders (VAEs) and deep metric learning to perform Bayesian optimisation (BO) over high-dimensional and structured input spaces. By…