activity
20182022
most citedSMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving

103 citations · 273 across the 23 of their papers we have counts for

collaborators
Showing cs.LGShow all

15 papers · 1 filter

cs.LG2022

Sample-Efficient Optimisation with Probabilistic Transformer Surrogates

Alexandre Maraval, Matthieu Zimmer, Antoine Grosnit +3

Faced with problems of increasing complexity, recent research in Bayesian Optimisation (BO) has focused on adapting deep probabilistic models as flexible alternatives to Gaussian P…

cs.LG2022

Reinforcement Learning in Presence of Discrete Markovian Context Evolution

Hang Ren, Aivar Sootla, Taher Jafferjee +3

We consider a context-dependent Reinforcement Learning (RL) setting, which is characterized by: a) an unknown finite number of not directly observable contexts; b) abrupt (disconti…

cs.LG2022

Learning to Identify Top Elo Ratings: A Dueling Bandits Approach

Xue Yan, Yali Du, Binxin Ru +3

The Elo rating system is widely adopted to evaluate the skills of (chess) game and sports players. Recently it has been also integrated into machine learning algorithms in evaluati…

cs.LG20212 cited

Revisiting the Characteristics of Stochastic Gradient Noise and Dynamics

Yixin Wu, Rui Luo, Chen Zhang +2

In this paper, we characterize the noise of stochastic gradients and analyze the noise-induced dynamics during training deep neural networks by gradient-based optimizers. Specifica…

cs.LG202111 cited

High-Dimensional Bayesian Optimisation with Variational Autoencoders and Deep Metric Learning

Antoine Grosnit, Rasul Tutunov, Alexandre Max Maraval +9

We introduce a method combining variational autoencoders (VAEs) and deep metric learning to perform Bayesian optimisation (BO) over high-dimensional and structured input spaces. By…

cs.LG2021

Efficient Semi-Implicit Variational Inference

Vincent Moens, Hang Ren, Alexandre Maraval +3

In this paper, we propose CI-VI an efficient and scalable solver for semi-implicit variational inference (SIVI). Our method, first, maps SIVI's evidence lower bound (ELBO) to a for…