5 papers · 1 filter
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
Lars van der Laan, Nathan Kallus
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate…
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
Lars van der Laan, Nathan Kallus
Fitted -iteration (FQI) and soft FQI are widely used value-based methods for offline reinforcement learning, but their standard stability guarantees often depend on Bellman comp…
Fitted Evaluation Without Bellman Completeness via Stationary Weighting
Lars van der Laan, Nathan Kallus
Fitted -evaluation (FQE) is a standard regression-based tool for off-policy evaluation, but existing stability guarantees often rely on Bellman completeness, a strong closure co…
Bellman Calibration for -Learning in Offline Reinforcement Learning
Lars van der Laan, Nathan Kallus
Reliable long-horizon value prediction is difficult in offline reinforcement learning because fitted value methods combine bootstrapping, function approximation, and distribution s…
Semiparametric Double Reinforcement Learning with Applications to Long-Term Causal Inference
Lars van der Laan, David Hubbard, Allen Tran +3
Double reinforcement learning (DRL) provides efficient off-policy inference for policy values in nonparametric Markov decision processes (MDPs), but fully nonparametric estimators…