5 papers
Refined Detection for Gumbel Watermarking
Tor Lattimore
We propose a simple detection mechanism for the Gumbel watermarking scheme proposed by Aaronson (2022). The new mechanism is proven to be near-optimal in a problem-dependent sense…
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
Tor Lattimore
We adapt the analysis of policy gradient for continuous time -armed stochastic bandits by Lattimore (2026) to the standard discrete time setup. As in continuous time, we prove t…
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
Tor Lattimore
We study a continuous-time diffusion approximation of policy gradient for -armed stochastic bandits. We prove that with a learning rate the regret is $O(k…
Bandit Convex Optimisation
Tor Lattimore
Bandit convex optimisation is a fundamental framework for studying zeroth-order convex optimisation. This book covers the many tools used for this problem, including cutting plane…
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
András György, Tor Lattimore, Nevena LaziÄ +1
Sound deductive reasoning -- the ability to derive new knowledge from existing facts and rules -- is an indisputably desirable aspect of general intelligence. Despite the major adv…