22 papers
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
Lars van der Laan, Nathan Kallus
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate…
Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach
Guang-Yuan Hao, Lars van der Laan, Aurélien Bibaut +1
We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in a new, different environmen…
Targeted maximum likelihood estimation of vaccine effectiveness and immune correlates in test-negative design studies with missing data
Leah I. B. Andrews, Lars van der Laan, Peter B. Gilbert
The test-negative design (TND) is a resource-efficient observational study design that can assess vaccine effectiveness and exposure-proximal immune correlates of disease. The TND…
Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration
Lars van der Laan, Nathan Kallus
Fitted -iteration (FQI) and soft FQI are widely used value-based methods for offline reinforcement learning, but their standard stability guarantees often depend on Bellman comp…
Fitted Evaluation Without Bellman Completeness via Stationary Weighting
Lars van der Laan, Nathan Kallus
Fitted -evaluation (FQE) is a standard regression-based tool for off-policy evaluation, but existing stability guarantees often rely on Bellman completeness, a strong closure co…
Bellman Calibration for -Learning in Offline Reinforcement Learning
Lars van der Laan, Nathan Kallus
Reliable long-horizon value prediction is difficult in offline reinforcement learning because fitted value methods combine bootstrapping, function approximation, and distribution s…