2 papers
cs.LG2023
Thompson Sampling for Linear Bandit Problems with Normal-Gamma Priors
Björn Lindenberg, Karl-Olof Lindahl
We consider Thompson sampling for linear bandit problems with finitely many independent arms, where rewards are sampled from normal distributions that are linearly dependent on unk…
cs.LG2021
Conjugated Discrete Distributions for Distributional Reinforcement Learning
Björn Lindenberg, Jonas Nordqvist, Karl-Olof Lindahl
In this work we continue to build upon recent advances in reinforcement learning for finite Markov processes. A common approach among previous existing algorithms, both single-acto…