153 citations · 362 across the 31 of their papers we have counts for
6 papers · 1 filter
Probabilistic Inference in Reinforcement Learning Done Right
Jean Tarbouriech, Tor Lattimore, Brendan O'Donoghue
A popular perspective in Reinforcement learning (RL) casts the problem as probabilistic inference on a graphical model of the Markov decision process (MDP). The core object of stud…
Context-lumpable stochastic bandits
Chung-Wei Lee, Qinghua Liu, Yasin Abbasi-Yadkori +3
We consider a contextual bandit problem with contexts and actions. In each round , the learner observes a random context and chooses an action based on its pas…
Sequential Best-Arm Identification with Application to Brain-Computer Interface
Xin Zhou, Botao Hao, Jian Kang +2
A brain-computer interface (BCI) is a technology that enables direct communication between the brain and an external device or computer system. It allows individuals to interact wi…
A Second-Order Method for Stochastic Bandit Convex Optimisation
Tor Lattimore, András György
We introduce a simple and efficient algorithm for unconstrained zeroth-order stochastic convex bandits and prove its regret is at most $(1 + r/d)[d^{1.5} \sqrt{n} + d^3] polylog(n,…
Linear Partial Monitoring for Sequential Decision-Making: Algorithms, Regret Bounds and Applications
Johannes Kirschner, Tor Lattimore, Andreas Krause
Partial monitoring is an expressive framework for sequential decision-making with an abundance of applications, including graph-structured and dueling bandits, dynamic pricing and…
Leveraging Demonstrations to Improve Online Learning: Quality Matters
Botao Hao, Rahul Jain, Tor Lattimore +2
We investigate the extent to which offline demonstration data can improve online learning. It is natural to expect some improvement, but the question is how, and by how much? We sh…