Showing cs.CEShow all
2 papers · 1 filter
cs.CE2026
Fitting Reinforcement Learning Model to Behavioral Data under Bandits
Hao Zhu, Jasper Hoffmann, Baohe Zhang +1
We consider the problem of fitting a reinforcement learning (RL) model to some given behavioral data under a multi-armed bandit environment. These models have received much attenti…
cs.CE2025
Solving Inverse Problem for Multi-armed Bandits via Convex Optimization
Hao Zhu, Joschka Boedecker
We consider the inverse problem of multi-armed bandits (IMAB) that are widely used in neuroscience and psychology research for behavior modelling. We first show that the IMAB probl…