activity
20242026
collaborators

22 papers

cs.LG2026

Foundations of Reinforcement Learning and Control:Connections and New Perspectives

Claire Vernade, Onno Eberhard, Martha White +4

Reinforcement learning and control theory are two adjacent scientific fields that focus on optimizing the controller of unknown dynamical systems using feedback. While both fields…

cs.LG2026

Learning to Reason Efficiently with Discounted Reinforcement Learning

Alex Ayoub, Kavosh Asadi, Dale Schuurmans +2

Large reasoning models (LRMs) often consume excessive tokens, inflating computational cost and latency. More broadly, in goal reaching sequential decision problems we often want to…

cs.LG2026

Randomized Exploration for Linear Bandits via Absolute Perturbations

Toshinori Kitamura, Shuai Liu, Csaba Szepesvári

In stochastic linear bandits, the canonical Upper Confidence Bound (UCB) algorithm admits a simple frequentist regret analysis but can be computationally demanding, while Thompson…

cs.LG2026

Sharp analysis of linear ensemble sampling

David Janz, Arya Akhavan, Csaba Szepesvári

We analyse linear ensemble sampling (ES) with standard Gaussian perturbations in stochastic linear bandits. We show that for ensemble size , ES attains $\tilde O(d^{…

cs.LG2026

Exploration via linearly perturbed loss minimisation

David Janz, Shuai Liu, Alex Ayoub +1

We introduce exploration via linear loss perturbations (EVILL), a randomised exploration method for structured stochastic bandit problems that works by solving for the minimiser of…

cs.LG2026

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear -Realizability and Concentrability

Volodymyr Tkachuk, Csaba Szepesvári, Xiaoqi Tan

We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization. Prior work established that statisticall…