activity
20192023
most citedExploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models

4 citations · 10 across the 3 of their papers we have counts for

collaborators

6 papers

q-fin.PR2023★ 3 cited

Insurance pricing on price comparison websites via reinforcement learning

Tanut Treetanthiploet, Yufei Zhang, Lukasz Szpruch +4

The emergence of price comparison websites (PCWs) has presented insurers with unique challenges in formulating effective pricing strategies. Operating on PCWs requires insurers to…

cs.LG2022★ 3 cited

Optimal scheduling of entropy regulariser for continuous-time linear-quadratic reinforcement learning

Lukasz Szpruch, Tanut Treetanthiploet, Yufei Zhang

This work uses the entropy-regularised relaxed stochastic control perspective as a principled framework for designing reinforcement learning (RL) algorithms. Herein agent interacts…

cs.LG2021★ 4 cited

Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models

Lukasz Szpruch, Tanut Treetanthiploet, Yufei Zhang

We develop a probabilistic framework for analysing model-based reinforcement learning in the episodic setting. We then apply it to study finite-time horizon stochastic control prob…

math.OC2021

Generalised correlated batched bandits via the ARC algorithm with application to dynamic pricing

Samuel Cohen, Tanut Treetanthiploet

The Asymptotic Randomised Control (ARC) algorithm provides a rigorous approximation to the optimal strategy for a wide class of Bayesian bandits, while retaining low computational…

math.OC2020

Asymptotic Randomised Control with applications to bandits

Samuel N. Cohen, Tanut Treetanthiploet

We consider a general multi-armed bandit problem with correlated (and simple contextual and restless) elements, as a relaxed control problem. By introducing an entropy regularisati…

math.OC2019

Gittins' theorem under uncertainty

Samuel N. Cohen, Tanut Treetanthiploet

We study dynamic allocation problems for discrete time multi-armed bandits under uncertainty, based on the the theory of nonlinear expectations. We show that, under strong independ…