3 papers
cs.LG2026
Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead Thresholding
Jiamin Xu, Kyra Gan
Online reinforcement learning in non-episodic, finite-horizon MDPs remains underexplored and is challenged by the need to estimate returns to a fixed terminal time. Existing infini…
cs.LG2026
Integrating Causal DAGs in Deep RL: Activating Minimal Markovian States with Multi-Order Exposure
Jiamin Xu, Jacqueline Maasch, Kyra Gan
Online reinforcement learning (RL) relies on the Markov property for guaranteed performance, but real-world applications often lack well-defined states given raw observed variables…
cs.LG2026
From Restless to Contextual: A Thresholding Bandit Reformulation For Finite-horizon Improvement
Jiamin Xu, Ivan Nazarov, Aditya Rastogi +2
This paper addresses the poor finite-horizon performance of existing online \emph{restless bandit} (RB) algorithms, which stems from the prohibitive sample complexity of learning a…