2 papers
cs.LG2026
Planning Under Observation Mismatch for Traffic Signal Control via Adaptive Modular World Models
Zherui Huang, Yicheng Liu, Chumeng Liang +1
Deploying learned decision-making systems often requires transferring to new sites where the sensing pipeline differs. In such cases, observations can change in semantics and dimen…
math.OC2025
Fitted Q-Iteration via Max-Plus-Linear Approximation
Y. Liu, M. A. S. Kolarijani
In this study, we consider the application of max-plus-linear approximators for Q-function in offline reinforcement learning of discounted Markov decision processes. In particular,…