activity
20242026
collaborators

22 papers

stat.ML2026

Fitted Occupancy-Ratio Evaluation without Bellman Completeness

Lars van der Laan, Nathan Kallus

Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate…

cs.LG2026

Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach

Guang-Yuan Hao, Lars van der Laan, Aurélien Bibaut +1

We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in a new, different environmen…

stat.ME2026

Targeted maximum likelihood estimation of vaccine effectiveness and immune correlates in test-negative design studies with missing data

Leah I. B. Andrews, Lars van der Laan, Peter B. Gilbert

The test-negative design (TND) is a resource-efficient observational study design that can assess vaccine effectiveness and exposure-proximal immune correlates of disease. The TND…

stat.ML2026

Stationary Reweighting Yields Local Convergence of Soft Fitted Q-Iteration

Lars van der Laan, Nathan Kallus

Fitted -iteration (FQI) and soft FQI are widely used value-based methods for offline reinforcement learning, but their standard stability guarantees often depend on Bellman comp…

stat.ML2026

Fitted Evaluation Without Bellman Completeness via Stationary Weighting

Lars van der Laan, Nathan Kallus

Fitted -evaluation (FQE) is a standard regression-based tool for off-policy evaluation, but existing stability guarantees often rely on Bellman completeness, a strong closure co…

stat.ML2026

Bellman Calibration for -Learning in Offline Reinforcement Learning

Lars van der Laan, Nathan Kallus

Reliable long-horizon value prediction is difficult in offline reinforcement learning because fitted value methods combine bootstrapping, function approximation, and distribution s…