11 papers
Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning
Lars van der Laan, Nathan Kallus
Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equ…
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
Lars van der Laan, Nathan Kallus
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate…
Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach
Guang-Yuan Hao, Lars van der Laan, Aurélien Bibaut +1
We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in a new, different environmen…
Efficient Inference for Inverse Reinforcement Learning and Dynamic Discrete Choice Models
Lars van der Laan, Aurélien Bibaut, Nathan Kallus
In many sequential decision-making problems, researchers observe actions but not the rewards that drive behavior, yet still wish to evaluate and compare counterfactual policies. In…
Soft Fitted Q-Iteration without Bellman Completeness: Occupancy Reweighting and Temperature Annealing
Lars van der Laan, Nathan Kallus
Fitted \(Q\)-iteration (FQI) is a standard regression-based method for optimal control in offline reinforcement learning, but its stability under function approximation often relie…
Fitted Q-Evaluation without Bellman Completeness via Occupancy Weighting
Lars van der Laan, Nathan Kallus
Fitted \(Q\)-evaluation (FQE) is a standard regression-based method for off-policy evaluation, but under distribution shift, value-function realizability alone does not ensure conv…