collaborators

7 papers

cs.LG2026

Probabilistic Recurrent Intention Switching Model

Wenyuan Sheng, Hao Zhu, Joschka Boedecker

Inverse reinforcement learning (IRL) recovers reward functions from observed behavior, yet traditional methods assume a single stationary reward that cannot capture goal switching…

cs.LG2026

Spectral Alignment in Forward-Backward Representations via Temporal Abstraction

Seyed Mahdi B. Azad, Jasper Hoffmann, Iman Nematollahi +3

Forward-backward (FB) representations provide a powerful framework for learning the successor representation (SR) in continuous spaces by enforcing a low-rank factorization. Howeve…

cs.CE2026

Fitting Reinforcement Learning Model to Behavioral Data under Bandits

Hao Zhu, Jasper Hoffmann, Baohe Zhang +1

We consider the problem of fitting a reinforcement learning (RL) model to some given behavioral data under a multi-armed bandit environment. These models have received much attenti…

math.OC2025

Disciplined Biconvex Programming

Hao Zhu, Joschka Boedecker

We introduce disciplined biconvex programming (DBCP), a modeling framework for specifying and solving biconvex optimization problems. Biconvex optimization problems arise in variou…

math.OC2025

Multi-convex Programming for Discrete Latent Factor Models Prototyping

Hao Zhu, Shengchao Yan, Jasper Hoffmann +1

Discrete latent factor models (DLFMs) are widely used in various domains such as machine learning, economics, neuroscience, psychology, etc. Currently, fitting a DLFM to some datas…

cs.CE2025

Solving Inverse Problem for Multi-armed Bandits via Convex Optimization

Hao Zhu, Joschka Boedecker

We consider the inverse problem of multi-armed bandits (IMAB) that are widely used in neuroscience and psychology research for behavior modelling. We first show that the IMAB probl…