2 papers
cs.LG2019
Emergent Tool Use From Multi-Agent Autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov +4
Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocu…
cs.LG2018
Best arm identification in multi-armed bandits with delayed feedback
Aditya Grover, Todor Markov, Peter Attia +8
We propose a generalization of the best arm identification problem in stochastic multi-armed bandits (MAB) to the setting where every pull of an arm is associated with delayed feed…