4 papers
Cascaded consensus splitting for multi-branch contingency games
Bastien Lechardoy, Pau de las Heras Molins, Thibault Lahire +4
Contingency games enable agents to anticipate and plan for other agents' hypothetical intents by constructing trajectories with a shared prefix and intent-dependent branches. While…
Importance Sampling for Stochastic Gradient Descent in Deep Neural Networks
Thibault Lahire
Stochastic gradient descent samples uniformly the training set to build an unbiased gradient estimate with a limited number of samples. However, at a given step of the training pro…
Actor Loss of Soft Actor Critic Explained
Thibault Lahire
This technical report is devoted to explaining how the actor loss of soft actor critic is obtained, as well as the associated gradient estimate. It gives the necessary mathematical…
Large Batch Experience Replay
Thibault Lahire, Matthieu Geist, Emmanuel Rachelson
Several algorithms have been proposed to sample non-uniformly the replay buffer of deep Reinforcement Learning (RL) agents to speed-up learning, but very few theoretical foundation…