activity
20162026
most citedDegenerate Feedback Loops in Recommender Systems

153 citations · 362 across the 31 of their papers we have counts for

collaborators
Showing 2023Show all

6 papers · 1 filter

cs.LG2023

Probabilistic Inference in Reinforcement Learning Done Right

Jean Tarbouriech, Tor Lattimore, Brendan O'Donoghue

A popular perspective in Reinforcement learning (RL) casts the problem as probabilistic inference on a graphical model of the Markov decision process (MDP). The core object of stud…

cs.LG2023

Context-lumpable stochastic bandits

Chung-Wei Lee, Qinghua Liu, Yasin Abbasi-Yadkori +3

We consider a contextual bandit problem with contexts and actions. In each round , the learner observes a random context and chooses an action based on its pas…

cs.HC20232 cited

Sequential Best-Arm Identification with Application to Brain-Computer Interface

Xin Zhou, Botao Hao, Jian Kang +2

A brain-computer interface (BCI) is a technology that enables direct communication between the brain and an external device or computer system. It allows individuals to interact wi…

cs.LG2023

A Second-Order Method for Stochastic Bandit Convex Optimisation

Tor Lattimore, András György

We introduce a simple and efficient algorithm for unconstrained zeroth-order stochastic convex bandits and prove its regret is at most $(1 + r/d)[d^{1.5} \sqrt{n} + d^3] polylog(n,…

cs.LG2023

Linear Partial Monitoring for Sequential Decision-Making: Algorithms, Regret Bounds and Applications

Johannes Kirschner, Tor Lattimore, Andreas Krause

Partial monitoring is an expressive framework for sequential decision-making with an abundance of applications, including graph-structured and dueling bandits, dynamic pricing and…

cs.LG2023

Leveraging Demonstrations to Improve Online Learning: Quality Matters

Botao Hao, Rahul Jain, Tor Lattimore +2

We investigate the extent to which offline demonstration data can improve online learning. It is natural to expect some improvement, but the question is how, and by how much? We sh…